This is an automated email from the ASF dual-hosted git repository.
kparisa pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/iggy-website.git
The following commit(s) were added to refs/heads/main by this push:
new b4c00436 chore(docs): add server tuning & Linux optimization (#88)
b4c00436 is described below
commit b4c0043664761d26e95f137ce0ad71ccfd621c37
Author: Piotr Gankiewicz <[email protected]>
AuthorDate: Wed Sep 9 10:37:36 2026 +0200
chore(docs): add server tuning & Linux optimization (#88)
Co-authored-by: hubcio <[email protected]>
Co-authored-by: Kranti Parisa <[email protected]>
---
content/docs/binary-protocol/commands.mdx | 6 +-
content/docs/binary-protocol/encodings.mdx | 2 +-
content/docs/clustering/deploy.mdx | 8 +-
content/docs/clustering/durability.mdx | 137 +++++++++++++++++++++++
content/docs/clustering/meta.json | 2 +-
content/docs/introduction/architecture.mdx | 12 ++-
content/docs/introduction/concepts.mdx | 2 +-
content/docs/introduction/getting-started.mdx | 4 +-
content/docs/server/benchmarking.mdx | 34 ++++--
content/docs/server/configuration.mdx | 82 +++++++-------
content/docs/server/docker.mdx | 12 +--
content/docs/server/durability.mdx | 150 ++++++++++++++++++++++++++
content/docs/server/linux-tuning.mdx | 145 +++++++++++++++++++++++++
content/docs/server/meta.json | 2 +-
content/docs/server/security.mdx | 4 +-
content/docs/server/storage-engine.mdx | 69 ++++++------
content/docs/server/topic-options.mdx | 22 ++--
17 files changed, 572 insertions(+), 121 deletions(-)
diff --git a/content/docs/binary-protocol/commands.mdx
b/content/docs/binary-protocol/commands.mdx
index 52580503..c97ec88c 100644
--- a/content/docs/binary-protocol/commands.mdx
+++ b/content/docs/binary-protocol/commands.mdx
@@ -80,7 +80,7 @@ LEAVE_CONSUMER_GROUP = 605 # operation 149
SYNC_CONSUMER_GROUP = 606 # non-replicated
```
-`FLUSH_UNSAVED_BUFFER` still decodes on the wire, but the server has no
on-demand flush primitive and answers every call with `FeatureUnavailable`.
Per-topic durability is configured with the `enforce_fsync` [topic
option](/docs/server/topic-options) instead.
+`FLUSH_UNSAVED_BUFFER` still decodes on the wire, but the server has no
on-demand flush primitive and answers every call with `FeatureUnavailable`.
Per-topic message completion is configured with the `durability` [topic
option](/docs/server/topic-options).
## Payloads
@@ -178,7 +178,7 @@ Patch semantics: absent option keys are left unchanged.
[stream_id: Identifier][partitions_count: u32][name_len: u8][name: N][options
block to end]
```
-The fixed fields are the shape of the operation: which stream, how many
partitions, what name. Every topic setting (`compression_algorithm`,
`message_expiry`, `max_topic_size`, `segment_size`, `enforce_fsync`,
`messages_required_to_save`, `size_of_messages_required_to_save`,
`preallocate_segments`) rides the options block. See [Topic
options](/docs/server/topic-options). `partitions_count` is an argument, not a
setting: it is consumed at admission and never persisted as an option.
+The fixed fields are the shape of the operation: which stream, how many
partitions, what name. Every topic setting (`compression_algorithm`,
`message_expiry`, `max_topic_size`, `segment_size`, `durability`,
`consumer_offset_durability`, `messages_required_to_save`,
`size_of_messages_required_to_save`, `preallocate_segments`) rides the options
block. See [Topic options](/docs/server/topic-options). `partitions_count` is
an argument, not a setting: it is consumed at admission and never per [...]
**Delete topic. Code: 303.**
@@ -262,7 +262,7 @@ The 16-byte prefix is followed by a stream of [batch
records](/docs/binary-proto
[stream_id: Identifier][topic_id: Identifier][partition_id: u32][fsync: u8]
```
-Parses, but the server always answers `FeatureUnavailable`: there is no
on-demand flush primitive. Use the per-topic `enforce_fsync` option for
durability guarantees.
+Parses, but the server always answers `FeatureUnavailable`: there is no
on-demand flush primitive. Select `durability=persisted` at topic creation to
require recoverable stable-storage copies before successful message completion.
### Consumer offsets
diff --git a/content/docs/binary-protocol/encodings.mdx
b/content/docs/binary-protocol/encodings.mdx
index 7f482db4..515e46f2 100644
--- a/content/docs/binary-protocol/encodings.mdx
+++ b/content/docs/binary-protocol/encodings.mdx
@@ -114,7 +114,7 @@ Semantics:
- **Update** requests are patches: keys absent from the block are left alone,
never reset. A client built before a key existed cannot erase it.
- Unknown keys are rejected at the wire edge, never silently skipped.
-The catalog is discoverable at runtime with `DESCRIBE_OPTIONS` (code 13),
payload `[scope: u8]` with scope `1` = topic, `2` = stream, `3` = user (HTTP:
`GET /options/topic`). Today only topics have keys. The stream and user
catalogs are empty, and any key sent for them is rejected. The topic catalog
(`segment_size`, `enforce_fsync`, `message_expiry`, `max_topic_size`, and the
rest) with defaults and constraints is documented on the [Topic
options](/docs/server/topic-options) page. `Updat [...]
+The catalog is discoverable at runtime with `DESCRIBE_OPTIONS` (code 13),
payload `[scope: u8]` with scope `1` = topic, `2` = stream, `3` = user (HTTP:
`GET /options/topic`). Today only topics have keys. The stream and user
catalogs are empty, and any key sent for them is rejected. The topic catalog
(`segment_size`, `durability`, `consumer_offset_durability`, `message_expiry`,
`max_topic_size`, and the rest) with defaults and constraints is documented on
the [Topic options](/docs/server/ [...]
## Compression
diff --git a/content/docs/clustering/deploy.mdx
b/content/docs/clustering/deploy.mdx
index 507a2902..72d0d8c1 100644
--- a/content/docs/clustering/deploy.mdx
+++ b/content/docs/clustering/deploy.mdx
@@ -85,13 +85,13 @@ Start each replica in a separate terminal. Use a different
data path for every p
```bash
# Replica 0
-IGGY_SYSTEM_PATH=local_data/node-0 ./target/debug/iggy-server --replica-id 0
+IGGY_PATH=local_data/node-0 ./target/debug/iggy-server --replica-id 0
# Replica 1
-IGGY_SYSTEM_PATH=local_data/node-1 ./target/debug/iggy-server --replica-id 1
+IGGY_PATH=local_data/node-1 ./target/debug/iggy-server --replica-id 1
# Replica 2
-IGGY_SYSTEM_PATH=local_data/node-2 ./target/debug/iggy-server --replica-id 2
+IGGY_PATH=local_data/node-2 ./target/debug/iggy-server --replica-id 2
```
The exported settings must be present in all three terminals.
@@ -103,7 +103,7 @@ In cluster mode:
- `ports` is the single source of listener ports: every enabled transport
needs an explicit per-node port, otherwise the server **refuses to start**
- `tcp_replica` carries replica-to-replica consensus traffic and is **always
required**
- `ip` must be a **literal IP address**. Use `advertised_address` when clients
can't reach it (see [Configuration](/docs/clustering/configuration))
-- use a different `system.path` for each process on the same host
+- use a different root `path` for each process on the same host
## Spanning multiple hosts
diff --git a/content/docs/clustering/durability.mdx
b/content/docs/clustering/durability.mdx
new file mode 100644
index 00000000..ec1c261e
--- /dev/null
+++ b/content/docs/clustering/durability.mdx
@@ -0,0 +1,137 @@
+---
+title: Cluster Durability
+description: "Quorum completion, persisted prepare history, and the failures
each topic policy covers."
+---
+
+Iggy uses [Viewstamped Replication](/docs/clustering/vsr) to commit partition
+operations. The topic's `durability` and `consumer_offset_durability`
+options select the storage guarantee required for message production and
+explicit offset changes, respectively. Both default independently to
+`replicated`. Both continue to write data to disk.
+
+For the creation options and single-node behavior, see
+[Durability](/docs/server/durability).
+
+## What an acknowledgement means
+
+For an operation that waits for completion:
+
+| Policy | Required before success |
+| --- | --- |
+| `replicated` | VSR quorum commit and local application, without an
additional stable-storage barrier |
+| `persisted` | VSR quorum commit backed by recoverable stable-storage copies
at the required quorum, followed by local application |
+
+The policy changes what a replica must retain before its acknowledgement can
+count toward commit. It does not replace quorum commit with a primary-only
+disk write.
+
+Iggy's replication quorum is not a strict majority for every group size:
+
+| Replicas | Replication quorum | View-change quorum |
+| --- | --- | --- |
+| 1 | 1 | 1 |
+| 2 | 2 | 2 |
+| 3 | 2 | 2 |
+| 4 | 2 | 3 |
+| 5 | 3 | 3 |
+| 6 | 3 | 4 |
+
+Except for two replicas, the replication quorum is `min(ceil(n / 2), 3)`
+and the view-change quorum is `n - replication_quorum + 1`. Two replicas
+require both for either quorum. The quorums therefore intersect. These are
+the [implemented quorum
rules](https://github.com/apache/iggy/blob/97f7b0c0335f81691bef34f923f21d783016a6a4/core/consensus/src/impls.rs),
+not a configurable acknowledgement count.
+
+## Replicated completion
+
+With both policies set to `replicated`, partition prepares stay off the
+disk prepare-WAL path. Replicas retain the prepares in memory, and committed
+messages reach segment files through ordinary flush scheduling.
+
+Success therefore does not prove that the operation reached stable storage
+on any replica. Replication protects against failures while sufficient peers
+retain the history and can form the quorums needed for recovery and progress.
+For example, a healthy three-replica group can continue after one replica
+fails.
+
+Failure independence matters. A shared power loss, an OOM cascade, or a
+crash triggered on every replica can destroy multiple volatile copies at
+once. A process-only failure can lose messages still in the processes'
+journals, even while the kernels and their page caches remain alive.
+Replication count alone does not establish a bound on that loss.
+
+## Persisted completion
+
+When either policy is `persisted`, each multi-replica partition uses a
+bounded on-disk prepare WAL. The WAL records full prepares, including message
+payloads. A prepare requiring persistence cannot release its `PrepareOk`
+until its history and durable frontier are recoverable.
+
+Prepares can be forwarded while local persistence is pending. Once enough
+replicas have met the required barrier, the operation can commit and the
+primary can apply it and reply. An acknowledged message may still be in
+the prepare WAL rather than in a segment file; its recovery does not depend
+on first reaching a segment flush threshold.
+
+The WAL also retains predecessors across message and offset operations.
+Consequently:
+
+- `durability=persisted` does not make an offset response persisted when
+ `consumer_offset_durability=replicated`.
+- `consumer_offset_durability=persisted` with replicated messages still
+ journals message payloads. A durable offset's predecessor history must
+ remain recoverable.
+- A shared barrier can also persist co-batched operations with the weaker
+ policy. That incidental persistence does not strengthen what their earlier
+ `replicated` acknowledgements promised.
+- Only the `replicated` / `replicated` combination avoids the disk prepare
+ WAL entirely.
+
+WAL history is reclaimed only after the materialized segment and offset
+state needed to replace it has been synchronized. The server's
+`[partition] wal_bytes_max` setting, default `256 MiB`, bounds active WAL
+and queued/in-flight prepare bytes per partition. Capacity pressure causes
+checkpointing and backpressure; it does not downgrade a persisted operation.
+Temporary rewrites need additional disk space.
+
+## Recovery and failure limits
+
+A restarting replica loads segment and offset state and, when enabled,
+reconciles the durable prepare WAL with that state. It also restores durable
+consensus state and follows the recovery protocol before serving as a healthy
+replica. It does not treat every locally recovered prepare as committed.
+
+Missing required history, checksum failures, or contradictory materialized
+state must not become an empty healthy partition. Recovery can fence a
+partition and repair it from peers. Storage errors withhold successful
+completion; a locally failed application of a committed operation fences the
+partition and initiates server shutdown.
+
+`persisted` protects acknowledged operations across process and power
+failures when the required storage copies remain intact and the storage
+honors synchronization. Recovery and availability still require the
+protocol's quorums. It cannot protect against destruction of all durable
+copies, nor does it prevent configured retention or explicit deletion.
+
+## Choosing the policies
+
+Use `replicated` when the workload accepts the risk of losing acknowledged
+operations after correlated failures in exchange for avoiding a required
+storage barrier at completion. Isolate replicas across failure domains; see
+[Deployment](/docs/clustering/deploy).
+
+Use `durability=persisted` when acknowledged messages must be recoverable
+after volatile copies are lost. Select
+`consumer_offset_durability=persisted` independently when explicit offset
+stores and deletes need the same guarantee. Ordinary flush thresholds can
+remain at their defaults.
+
+Neither policy turns HTTP `ack=none` or a poll's auto-commit into an awaited
+durable result. See [Which responses prove
completion](/docs/server/durability#which-responses-prove-completion).
+
+The storage mechanisms are implemented in the
+[partition persistence
worker](https://github.com/apache/iggy/blob/97f7b0c0335f81691bef34f923f21d783016a6a4/core/partitions/src/persistence.rs)
+and [prepare
journal](https://github.com/apache/iggy/blob/97f7b0c0335f81691bef34f923f21d783016a6a4/core/journal/src/partition_journal.rs).
+The [crash-recovery
tests](https://github.com/apache/iggy/blob/97f7b0c0335f81691bef34f923f21d783016a6a4/core/integration/tests/cluster/crash_durability.rs)
+exercise acknowledged persisted messages and offsets below ordinary flush
+thresholds, on both a singleton and a three-replica cluster.
diff --git a/content/docs/clustering/meta.json
b/content/docs/clustering/meta.json
index 9d6ae8e6..fd86ff6a 100644
--- a/content/docs/clustering/meta.json
+++ b/content/docs/clustering/meta.json
@@ -1,4 +1,4 @@
{
"title": "Clustering",
- "pages": ["vsr", "deploy", "configuration", "security", "client-failover"]
+ "pages": ["vsr", "durability", "deploy", "configuration", "security",
"client-failover"]
}
diff --git a/content/docs/introduction/architecture.mdx
b/content/docs/introduction/architecture.mdx
index 3dbe2810..f6032bca 100644
--- a/content/docs/introduction/architecture.mdx
+++ b/content/docs/introduction/architecture.mdx
@@ -118,8 +118,10 @@ Iggy uses a custom memory pool with 28 buckets holding
buffer sizes from 4 KiB t
Messages flow through a multi-stage write pipeline:
-1. Messages arrive on the owning shard and are buffered in the partition
journal
-2. A flush is triggered when either the message count threshold or the size
threshold is reached - both are **per-topic options** set at topic creation
(defaults: 1024 messages, 1 MiB)
-3. The `MessagesWriter` uses **vectored I/O** with up to 1024 buffers per
syscall
-4. `fsync` per write is a per-topic option (`enforce_fsync`) for durability
guarantees
-5. When a segment reaches the topic's segment size (default 1 GiB), it is
**sealed** and a new segment is created
+1. Messages arrive on the owning shard and are buffered in the partition
journal.
+2. Partition VSR replicates prepares. With `durability=persisted`, a
multi-replica group requires recoverable prepare-WAL copies at the replication
quorum before commit.
+3. Committed operations are applied before success is returned. A singleton
with `durability=persisted` synchronizes local segment state before replying.
+4. Ordinary segment writes use per-topic count and byte thresholds (defaults:
1024 messages, 1 MiB); required persistence, capacity pressure, and lifecycle
work can flush earlier. The `MessagesWriter` uses **vectored I/O** with up to
1024 buffers per syscall.
+5. When a segment reaches the topic's segment size (default 1 GiB), it is
**sealed** and a new segment is created.
+
+Message `durability` and `consumer_offset_durability` default independently to
`replicated`. Both policies write data to disk; `persisted` adds a
stable-storage requirement at completion. See
[Durability](/docs/server/durability).
diff --git a/content/docs/introduction/concepts.mdx
b/content/docs/introduction/concepts.mdx
index 70d4af2f..e1157cd1 100644
--- a/content/docs/introduction/concepts.mdx
+++ b/content/docs/introduction/concepts.mdx
@@ -36,7 +36,7 @@ The stream is a logical concept, and you might think of it as
a **namespace**. F
The topic is also the logical concept, which is a part of the stream. The
topic is identified by its unique ID. You could think of topic as an entity
being responsible for storing the specific type of the records. For example,
you could have a topic for the user events, and another topic for the order
events, etc.
-The messages are not being stored in the topic directly, but rather in the
**partitions**, which are assigned to the topic. The topic can have one or more
partitions assigned, that could help achieve higher parallelism and throughput.
The topic can also have the **retention policy** assigned, which means that the
records are being deleted automatically once they are older than the specified
retention period. Topics also support maximum size limits and per-topic
durability options (`segme [...]
+The messages are not being stored in the topic directly, but rather in the
**partitions**, which are assigned to the topic. The topic can have one or more
partitions assigned, that could help achieve higher parallelism and throughput.
The topic can also have the **retention policy** assigned, which means that the
records are being deleted automatically once they are older than the specified
retention period. Topics also support maximum size limits and per-topic storage
options (`segment_ [...]
## Partition
diff --git a/content/docs/introduction/getting-started.mdx
b/content/docs/introduction/getting-started.mdx
index 97ba8c9e..1f5bc12f 100644
--- a/content/docs/introduction/getting-started.mdx
+++ b/content/docs/introduction/getting-started.mdx
@@ -33,7 +33,7 @@ docker run --rm \
apache/iggy:latest
```
-`SYS_NICE`, the seccomp setting and the memlock limit are all required by
`io_uring` and the thread-per-core architecture; see [Docker &
Helm](/docs/server/docker) for the details. `IGGY_TCP_ADDRESS` is needed
because the server binds to `127.0.0.1` inside the container by default, which
a published port cannot reach. `IGGY_NODE_ADVERTISED_ADDRESS` is needed because
that wildcard leaves the server with no address to give clients, and it refuses
to start rather than publish one nobody can [...]
+The capabilities, seccomp setting, and memlock limit form a permissive
development setup. Production deployments can use narrower syscall permissions
and a finite memory budget; see [Docker &
Helm](/docs/server/docker#why-these-capabilities) for the details.
`IGGY_TCP_ADDRESS` is needed because the server binds to `127.0.0.1` inside the
container by default, which a published port cannot reach.
`IGGY_NODE_ADVERTISED_ADDRESS` is needed because that wildcard leaves the
server with no addre [...]
Alternatively, build from source by cloning the
[repository](https://github.com/apache/iggy) and running:
@@ -201,7 +201,7 @@ async fn init_system(client: &IggyClient) {
}
```
-Every field of `TopicCreateOptions` left as `None` resolves to the **server
defaults**. It's also where the per-topic durability knobs live:
`segment_size`, `enforce_fsync` and the flush thresholds
(`messages_required_to_save`, `size_of_messages_required_to_save`) can all be
set at topic creation.
+Optional fields of `TopicCreateOptions` left as `None` resolve to the **server
defaults**. The `durability` and `consumer_offset_durability` fields use
`Durability::Replicated` or `Durability::Persisted` and default independently
to `Replicated`. They select the completion guarantee for messages and explicit
offset changes. Segment size and the flush thresholds
(`messages_required_to_save`, `size_of_messages_required_to_save`) are separate
creation options. See [Durability](/docs/server/ [...]
Finally, let's send some messages into our stream. We will implement the basic
loop with an interval between each iteration to simulate publishing the batch
of messages. Since the streaming server works directly with the binary data,
and couldn't care less about the (de)serialization format for the message
payload, it's really up to you, how to efficiently stream the messages for your
use case.
diff --git a/content/docs/server/benchmarking.mdx
b/content/docs/server/benchmarking.mdx
index a09ef288..65b4fe0f 100644
--- a/content/docs/server/benchmarking.mdx
+++ b/content/docs/server/benchmarking.mdx
@@ -21,24 +21,28 @@ First build the project in release mode:
cargo build --release
```
-**`iggy-bench` doesn't start a server.** It only connects, and fails with a
connection error when nothing is listening. Start `iggy-server` yourself first:
+**`iggy-bench` doesn't start a server.** It only connects, and fails with a
connection error when nothing is listening. For a disposable development server
with an empty data directory, use the credentials expected by the benchmark
defaults:
```bash
-cargo run --bin iggy-server -r
+cargo run --bin iggy-server -r -- --with-default-root-credentials
```
-or use the repo's driver script, which builds, starts the server, runs a send
and a poll benchmark, and shuts the server down:
+This flag affects only the first creation of the root user; existing
credentials remain unchanged, and environment variables take precedence. For
another server, pass its credentials with `iggy-bench -u USER -p PASSWORD`
before the benchmark kind.
+
+The repo's driver script builds, starts the server, runs a send and a poll
benchmark, and shuts the server down. **Use it only in a disposable checkout on
an isolated host: it deletes `local_data` and signals matching `iggy-server`
and `iggy-bench` processes by name.**
```bash
./scripts/run-benches.sh
```
-With a server running, invoke the benchmark kind and transport of your choice.
The tool creates the streams, topics, and partitions it needs, then sends or
polls messages.
+With a server running, invoke the benchmark kind and transport of your choice.
Workloads that produce create their benchmark streams and topics, deleting
pre-existing `bench-stream-{n}` streams unless `--reuse-streams` is set.
Consumer-only workloads (`pc` and `bcg`) require those streams and enough
messages to exist already. Run the corresponding producer workload first, with
matching stream, partition, and message counts.
## Benchmark kinds
Eight kinds are available, each with a short alias:
+Producer and consumer counts default to six. Pinned workloads also default to
six streams. Set explicit counts to match the workload and available CPUs.
+
1. Pinned producer (`pp`): N producers, each sending to its own stream-topic
with a single partition:
```bash
@@ -89,13 +93,13 @@ Eight kinds are available, each with a short alias:
## Transports
-Every kind runs over one of four transports: `tcp`, `quic`, `http`, or
`websocket` (alias `ws`). Each transport subcommand takes `--server-address`
for a non-default server, for example:
+The transport subcommands are `tcp`, `quic`, `http`, and `websocket` (alias
`ws`). Consumer-group workloads (`bcg`, `bpcg`, and `e2ecg`) require a binary
transport: HTTP consumer-group joining returns `FeatureUnavailable`, even
though the benchmark CLI accepts that combination. Each transport subcommand
takes `--server-address` for a non-default server, for example:
```bash
-cargo r --bin iggy-bench -r -- pinned-producer tcp --server-address
0.0.0.0:8090
+cargo r --bin iggy-bench -r -- pinned-producer tcp --server-address
127.0.0.1:8090
```
-Message count, batch size, message size, and the number of
producers/consumers/streams are all flags on the kind subcommand. `iggy-bench
--help` and `iggy-bench examples` show the full surface.
+Message count, batch size, message size, and both durability flags precede the
kind subcommand. Producer, consumer, stream, and partition counts belong to the
selected kind. `iggy-bench --help` and `iggy-bench examples` show the full
surface.
## Storing and comparing results
@@ -109,4 +113,18 @@ cargo r --bin iggy-bench -r -- pinned-producer tcp output
-o performance_results
## Performance
-The server is thread-per-core and shared-nothing, built on `io_uring` (via
`compio`), with shard and CPU pinning configurable under `[system.sharding]`.
Throughput and latency depend heavily on hardware, transport, and payload shape
(`messages-per-batch * message-size`). Run the benchmarks on your own hardware,
or browse current, dated results on the [benchmarking
platform](https://benchmarks.iggy.apache.org).
+The server is thread-per-core and shared-nothing, built on `io_uring` (via
`compio`), with shard and CPU pinning configurable under `[sharding]`.
Throughput and latency depend heavily on hardware, transport, and payload shape
(`messages-per-batch * message-size`). Run the benchmarks on your own hardware,
or browse current, dated results on the [benchmarking
platform](https://benchmarks.iggy.apache.org).
+
+## Prepare the host and topic policies
+
+Use the [Linux tuning guide](/docs/server/linux-tuning) to check process
limits, swap, huge pages, CPU/NUMA placement, writeback, and network capacity
before measuring. Record the exact host settings and change one variable at a
time. Use a separate load generator, or disjoint CPU sets when it shares a host
with the server.
+
+`--durability` and `--consumer-offset-durability` independently default to
`replicated`. Place them before the benchmark kind. For example, this keeps
consumer offsets replicated while measuring persisted message acknowledgments:
+
+```bash
+cargo r --bin iggy-bench -r -- --durability persisted
balanced-producer-and-consumer-group tcp
+```
+
+Add `--consumer-offset-durability persisted` to request that policy too. These
are create-time options and have no effect on existing topics with
`--reuse-streams`. Poll auto-commit remains asynchronous, so poll latency is
not an acknowledged offset-store measurement. `iggy-bench examples` shows all
benchmark kinds, transports, policy combinations, and available topic-option
examples.
+
+Keep both policies, replication-group size, batch and message sizes, CPU
allocation, and storage consistent between runs. Run long enough to include
flushes and checkpoints, and report tail latency, errors, memory pressure, and
client utilization alongside throughput.
diff --git a/content/docs/server/configuration.mdx
b/content/docs/server/configuration.mdx
index 8a863186..a5982da1 100644
--- a/content/docs/server/configuration.mdx
+++ b/content/docs/server/configuration.mdx
@@ -8,7 +8,6 @@ The server reads a single TOML file. A copy of
[`core/server/config.toml`](https
A config file doesn't have to be complete. It is **merged over the embedded
defaults**, so a minimal file overrides only what it names:
```toml
-[system]
path = "/var/lib/iggy"
[node]
@@ -23,6 +22,8 @@ address = "0.0.0.0:3000"
Both listeners here bind a wildcard, which tells the server nothing about
where clients reach it, so [`[node]`](#node) has to name that address
explicitly - without it this file is refused at boot.
+Host and kernel settings are configured separately from the server TOML. See
[Linux tuning](/docs/server/linux-tuning) for service limits, memory policies,
huge pages, CPU placement, and validation before deployment.
+
## How configuration loads
Configuration is resolved in three layers. Later layers win:
@@ -33,7 +34,7 @@ Configuration is resolved in three layers. Later layers win:
Before the environment is read, the server loads a `.env` file from the
working directory, or from the path named by `IGGY_ENV_PATH`.
-After boot the server writes the effective configuration, including the
addresses it actually bound, to `{system.path}/runtime/current_config.toml`.
+After boot the server writes the effective configuration, including the
addresses it actually bound, to `{path}/runtime/current_config.toml` with the
default runtime subdirectory.
### Environment variables
@@ -43,23 +44,23 @@ Every configuration key can be overridden with an `IGGY_`
variable. The name is
IGGY_TCP_ADDRESS=0.0.0.0:8090 # [tcp] address
IGGY_NODE_ADVERTISED_ADDRESS=iggy-1 # [node] advertised_address
IGGY_HTTP_ENABLED=true # [http] enabled
-IGGY_SYSTEM_PATH=/var/lib/iggy # [system] path
-IGGY_SYSTEM_LOGGING_LEVEL=debug # [system.logging] level
-IGGY_SYSTEM_SHARDING_CPU_ALLOCATION=4 # [system.sharding] cpu_allocation
+IGGY_PATH=/var/lib/iggy # root path
+IGGY_LOGGING_LEVEL=debug # [logging] level
+IGGY_SHARDING_CPU_ALLOCATION=4 # [sharding] cpu_allocation
```
Two variables live outside the config schema: `IGGY_ROOT_USERNAME` and
`IGGY_ROOT_PASSWORD` set the root credentials, always as a pair. **Only the
first creation** of the root user reads them. On an existing data directory the
stored root user is recovered unchanged.
### Secrets
-Five values are secret-flagged: `http.jwt.encoding_secret`,
`http.jwt.decoding_secret`, `system.encryption.key`,
`cluster.auth.shared_secret`, and `cluster.auth.previous_shared_secret`. Prefer
setting them through their environment variables
(`IGGY_HTTP_JWT_ENCODING_SECRET`, `IGGY_SYSTEM_ENCRYPTION_KEY`,
`IGGY_CLUSTER_AUTH_SHARED_SECRET`, and so on) instead of storing them in a
file. Secret values are masked in logs and **never serialized** into
`runtime/current_config.toml`.
+Five values are secret-flagged: `http.jwt.encoding_secret`,
`http.jwt.decoding_secret`, `encryption.key`, `cluster.auth.shared_secret`, and
`cluster.auth.previous_shared_secret`. Prefer setting them through their
environment variables (`IGGY_HTTP_JWT_ENCODING_SECRET`, `IGGY_ENCRYPTION_KEY`,
`IGGY_CLUSTER_AUTH_SHARED_SECRET`, and so on) instead of storing them in a
file. Secret values are masked in logs and **never serialized** into
`runtime/current_config.toml`.
### Validation at boot
-- Unknown keys are **silently ignored**. A typo in a key name leaves the
default in force, so check spelling when a setting doesn't seem to take effect.
The [relocated keys](#relocated-configuration-keys) below are the exception:
they refuse boot.
-- `system.segment.archive_expired = true` and
`system.recovery.recreate_missing_state = true` are unsupported placeholders.
Setting either aborts boot.
+- Unknown top-level TOML fields and unknown server `IGGY_` environment names
**reject startup**. Some nested tables can still ignore unknown fields, so
compare the effective configuration with the intended settings.
+- The old `[system]` table and its environment mappings are rejected. Removed
placeholders such as `archive_expired` and `recreate_missing_state` must be
removed, not set to `false`.
- Several sections validate relationships between keys (QUIC windows, sharding
shutdown budgets, metadata journal sizing, partition transfer floors). A
violation aborts boot with an error naming the keys. The constraints are listed
with their sections below.
-- The `RUST_LOG` environment variable **always takes precedence** over
`system.logging.level`.
+- The `RUST_LOG` environment variable **always takes precedence** over
`logging.level`.
## Command-line flags
@@ -67,27 +68,36 @@ Five values are secret-flagged: `http.jwt.encoding_secret`,
`http.jwt.decoding_s
| Flag | Description |
|------|-------------|
-| `--fresh`, `-f` | Delete the configured data directory (`local_data` by
default, see `IGGY_SYSTEM_PATH`) before boot and start on empty state. In
cluster mode this wipes **this replica only**; it rejoins and refills by state
transfer from the others. Wiping a quorum at the same time destroys committed
data. Do not put `--fresh` in a service unit: it would re-transfer the whole
dataset on every restart. |
+| `--fresh`, `-f` | Delete the configured data directory (`local_data` by
default, see `IGGY_PATH`) before boot and start on empty state. In cluster mode
this wipes **this replica only**; it rejoins and refills by state transfer from
the others. Wiping a quorum at the same time can destroy committed data. Do not
put `--fresh` in a service unit: it would re-transfer the whole dataset on
every restart. |
| `--with-default-root-credentials` | Set `IGGY_ROOT_USERNAME` and
`IGGY_ROOT_PASSWORD` to `iggy` unless they are already present in the
environment. Only the first creation of the root user reads these values.
Development only. |
| `--replica-id <N>` | Identify this node within `cluster.nodes`. Required
when `cluster.enabled = true`; the value must match exactly one `replica_id` in
the roster. |
## Relocated configuration keys
-Retention, durability, and segment layout are no longer server-wide settings.
They became per-topic options set at topic creation. See [Topic
options](/docs/server/topic-options). The server **refuses to boot** while any
of the old keys is still present, in the config file or in the environment, and
the error names the replacement. This is deliberate: the storage knobs are
create-only options now, so a topic created while an old key was silently
ignored could never be given the setting a [...]
+The server configuration no longer has a `[system]` wrapper. Its runtime
tables live at the root, and the data directory is the root `path` key. Old
`IGGY_SYSTEM_*` overrides are rejected. Retention, durability, and segment
layout are [topic creation options](/docs/server/topic-options), not
server-wide defaults.
| Old server key | Replacement |
|----------------|-------------|
+| `system.path` / `IGGY_SYSTEM_PATH` | Root `path` / `IGGY_PATH` |
+| `system.runtime`, `system.logging`, `system.encryption`,
`system.memory_pool`, `system.sharding` | Root tables `runtime`, `logging`,
`encryption`, `memory_pool`, `sharding`; remove `SYSTEM_` from their
environment names |
+| `system.partition.validate_checksum` | `partition.validate_checksum` |
+| `system.partition.path`, `partition.path` / `IGGY_PARTITION_PATH` | Removed.
The `partitions` path component is fixed. |
| `system.topic.max_size` | `max_topic_size` topic option |
| `system.topic.message_expiry` | `message_expiry` topic option |
-| `system.partition.enforce_fsync` | `enforce_fsync` topic option |
+| `system.partition.enforce_fsync` | `durability` topic option: `replicated`
or `persisted` |
+| `partition.consumer_offset_enforce_fsync` | `consumer_offset_durability`
topic option: `replicated` or `persisted` |
| `system.partition.messages_required_to_save` | `messages_required_to_save`
topic option |
| `system.partition.size_of_messages_required_to_save` |
`size_of_messages_required_to_save` topic option |
| `system.segment.size` | `segment_size` topic option |
| `system.segment.preallocate` | `preallocate_segments` topic option |
-| `system.message_deduplication` (whole section) | Removed with the feature.
No replacement. |
+| `system.stream.path`, `system.topic.path` | Removed. The `streams` and
`topics` path components are fixed; root `[stream]` and `[topic]` tables are
rejected. |
+| `system.segment.archive_expired`, `system.recovery.recreate_missing_state` |
Removed. Delete the keys and their old tables. |
+| `system.message_deduplication` (whole section) | Removed. No equivalent
configuration. |
| `extra` (whole section) | Removed. Admission caps are compile-time
constants. No replacement. |
-Other sections from older configs (`[message_saver]`, `[system.state]`,
`[system.backup]`, `[system.compression]`, `[tcp.socket]`, segment
`cache_indexes`, and similar) were removed without a boot check. Like any
unknown key they're silently ignored. Remove them from your file.
+Remove the old tables and environment variables after migrating their values.
Topic policies cannot be configured through TOML or `IGGY_` overrides, and the
former `enforce_fsync` topic option is also rejected. Both new durability
policies independently default to `replicated`; explicitly select `persisted`
at topic creation where required. Neither can be changed with `UpdateTopic`.
+
+Other removed sections, such as `[message_saver]`, `[system.state]`,
`[system.backup]`, `[system.compression]`, and `[tcp.socket]`, must also be
removed. Consult the embedded configuration for the deployed binary rather than
relying on ignored legacy fields.
## Reference
@@ -263,19 +273,19 @@ The `self_signed` flag doesn't mean the same thing
everywhere:
Ephemeral certificates **change on every start** and are meant for loopback
development, not production.
-### `[system]`
+### Root data path and runtime
| Key | Default | Description |
|-----|---------|-------------|
| `path` | `"local_data"` | Base directory for all server data. Everything
below is relative to it. |
-Storage sub-paths chain relative to each other: `[system.runtime] path =
"runtime"` (runtime data, including `current_config.toml`), `[system.logging]
path = "logs"`, `[system.stream] path = "streams"`, `[system.topic] path =
"topics"` (relative to the stream path), `[system.partition] path =
"partitions"` (relative to the topic path).
+Set `path` before any TOML table header. `[runtime] path = "runtime"` and
`[logging] path = "logs"` are relative to it. Data uses the fixed layout
`{path}/streams/{stream_id}/topics/{topic_id}/partitions/{partition_id}`.
Stream, topic, and partition path overrides are rejected.
-### `[system.logging]`
+### `[logging]`
| Key | Default | Description |
|-----|---------|-------------|
-| `path` | `"logs"` | Log directory, relative to `system.path`. |
+| `path` | `"logs"` | Log directory, relative to the root `path`. |
| `level` | `"info"` | Filter directive in `RUST_LOG` syntax: simple levels or
directives like `"warn,server=debug,iggy=trace"`. The `RUST_LOG` environment
variable always takes precedence. |
| `file_enabled` | `true` | Write logs to file as well as stdout. |
| `max_file_size` | `"500 MB"` | Size at which a log file rotates. `0` means
one unbounded file, which disables size-based rotation. |
@@ -283,32 +293,14 @@ Storage sub-paths chain relative to each other:
`[system.runtime] path = "runtim
| `rotation_check_interval` | `"1 h"` | How often rotation status is checked.
Avoid values below 1 s. |
| `retention` | `"7 days"` | How long log files are kept. Avoid values below 1
s. |
-### `[system.encryption]`
+### `[encryption]`
| Key | Default | Description |
|-----|---------|-------------|
-| `enabled` | `false` | Encrypt stored message payloads and state commands
with AES-256-GCM. |
-| `key` | `""` (empty) | 32-byte key, base64-encoded. Required when enabled.
Secret-flagged: prefer `IGGY_SYSTEM_ENCRYPTION_KEY`. |
-
-### `[system.partition]`
-
-| Key | Default | Description |
-|-----|---------|-------------|
-| `path` | `"partitions"` | Partition data directory, relative to the topic
path. |
-| `validate_checksum` | `true` | Re-hash every batch a disk poll reads and
fail the poll closed on a mismatch, so a segment damaged at rest is reported
instead of served. `false` serves whatever decodes, which can hand a consumer
bytes provably not the ones written. Only turn it off with a corruption guard
elsewhere in the stack. |
-
-Durability and flush cadence (`enforce_fsync`, `messages_required_to_save`,
`size_of_messages_required_to_save`) are per-topic creation options. See [Topic
options](/docs/server/topic-options).
+| `enabled` | `false` | Encrypt stored message payloads and user headers with
AES-256-GCM. Metadata journals, snapshots, and structural record headers remain
unencrypted. |
+| `key` | `""` (empty) | 32-byte key, base64-encoded. Required when enabled.
Secret-flagged: prefer `IGGY_ENCRYPTION_KEY`. |
-### `[system.segment]` and `[system.recovery]`
-
-| Key | Default | Description |
-|-----|---------|-------------|
-| `system.segment.archive_expired` | `false` | Unsupported placeholder.
Setting `true` aborts boot. |
-| `system.recovery.recreate_missing_state` | `false` | Unsupported
placeholder. Setting `true` aborts boot. |
-
-Segment size and preallocation are per-topic creation options (`segment_size`,
`preallocate_segments`).
-
-### `[system.memory_pool]`
+### `[memory_pool]`
| Key | Default | Description |
|-----|---------|-------------|
@@ -347,7 +339,7 @@ The Rust, Go, Python, Node, and async Java SDKs ping
automatically every 5 s, we
| `telemetry.traces.transport` | `"grpc"` | Trace export transport: `"grpc"`
or `"http"`. |
| `telemetry.traces.endpoint` | `"http://localhost:7281/v1/traces"` | Trace
export endpoint. |
-### `[system.sharding]`
+### `[sharding]`
| Key | Default | Description |
|-----|---------|-------------|
@@ -362,7 +354,7 @@ The Rust, Go, Python, Node, and async Java SDKs ping
automatically every 5 s, we
Valid `cpu_allocation` syntaxes:
- `"all"`: one shard per available CPU core
-- a number, e.g. `4`: four shards pinned to cores 0 to 3
+- a number, e.g. `4`: four shards on the first four CPUs in the process's
allowed set when `pin_cores` is enabled
- a range, e.g. `"5..8"`: three shards on cores 5, 6, 7
- `"numa:auto"`: all NUMA nodes and cores, NUMA-aware
- `"numa:nodes=0,1;cores=4;no_ht=true"`: NUMA nodes 0 and 1, four cores each,
no hyperthreads
@@ -379,11 +371,15 @@ Tunables for the metadata consensus plane (shard 0's VSR
replica: users, streams
### `[partition]`
-Per-partition consensus plane tunables. Unlike `[metadata]` (one plane on
shard 0), a pipeline exists per partition, so raising these **multiplies pinned
memory by the partition count**.
+Partition storage and consensus tunables share this table. Unlike `[metadata]`
(one plane on shard 0), a pipeline exists per partition, so raising
per-partition budgets **multiplies resource use by the partition count**.
+
+`durability`, `consumer_offset_durability`, segment size, preallocation, and
the message/byte flush thresholds are [topic creation
options](/docs/server/topic-options). `wal_bytes_max` sizes the persistence
machinery without selecting the topic's completion policy.
| Key | Default | Description |
|-----|---------|-------------|
| `prepare_queue_depth` | `32` | Uncommitted produce and consumer-offset ops
in flight per partition. Submits past it spill into a request queue of twice
this depth; once both are full the server drops the request without a reply and
the client retries on its own timeout. Must be between 1 and 127. |
+| `validate_checksum` | `true` | Re-hash batches read from segment storage and
report a mismatch instead of serving them. |
+| `wal_bytes_max` | `"256 MiB"` | Active WAL plus queued/in-flight prepare
budget per multi-replica partition when either topic policy is `persisted`. A 4
KiB multiple, from 128 MiB + 8 KiB through 4 GiB. Checkpointing reclaims
history only after materialized state is synchronized. Temporary rewrites
require extra disk space. Environment override: `IGGY_PARTITION_WAL_BYTES_MAX`.
|
| `evicted_ring_capacity` | `4096` | Entries retained per multi-replica
partition for journal repair after a peer rejoins. Must be between 1 and 65536.
Single-replica partitions retain nothing. |
| `evicted_ring_bytes_max` | `"16 MiB"` | Byte ceiling for the evicted ring;
whichever ring cap trips first evicts. At most `"256 MiB"`. |
| `transfer_served_cache_bytes_max` | `"2176 MiB"` | Byte budget, **per
shard**, for segment payloads kept resident to serve state-transfer chunk
requests. The default fits two sealed segments at the 1 GiB ceiling, each with
one max-message overshoot. Serving concurrency is `floor(this /
max(transfer_artifact_bytes_max, 1 GiB + 64 MiB))`, minimum one; boot warns
when it drops below two. At most `"64 GiB"`. |
diff --git a/content/docs/server/docker.mdx b/content/docs/server/docker.mdx
index d4df55bb..c46b89a3 100644
--- a/content/docs/server/docker.mdx
+++ b/content/docs/server/docker.mdx
@@ -7,15 +7,15 @@ description: "Run the Iggy server from the official Docker
images, and deploy it
You can easily run the Iggy server with Docker - the official images can be
found [here](https://hub.docker.com/r/apache/iggy), simply type `docker pull
apache/iggy`.
-Two properties of the published image matter for any deployment:
+These properties of the published image matter for deployment:
-- The working directory is `/app` and the `iggy-server` and `iggy` binaries
are on `PATH` (`/usr/local/bin`). The default data directory `local_data`
therefore resolves to `/app/local_data` - **mount your volume there**, or set
`IGGY_SYSTEM_PATH` and mount that path instead.
+- The working directory is `/app` and the `iggy-server` and `iggy` binaries
are on `PATH` (`/usr/local/bin`). The default data directory `local_data`
therefore resolves to `/app/local_data` - **mount your volume there**, or set
`IGGY_PATH` and mount that path instead.
- The image bakes in no address overrides, so the server binds the loopback
defaults (`127.0.0.1`) and is **unreachable from outside the container** even
with published ports. Set `IGGY_TCP_ADDRESS=0.0.0.0:8090` (and the equivalent
for every other transport you expose) alongside the `-p` flags.
- The wildcard says nothing about where clients reach the container, so the
server refuses to start until `IGGY_NODE_ADVERTISED_ADDRESS` supplies that
address. Use `localhost` when the ports are published to the host, the compose
service name when the clients are containers on the same network, and the
external hostname or load balancer name when they are further away. The value
reaches clients through cluster metadata, where it is the endpoint they
reconnect through.
Below is an example `docker-compose.yml` which overrides the default
configuration (see [Configuration](/docs/server/configuration)) with
environment variables. If you prefer using the configuration file, you can
mount it as a volume and provide the path to it with the `IGGY_CONFIG_PATH`
environment variable.
-When running the container, **make sure to include the additional
capabilities** required by `io_uring` and CPU affinity:
+The examples use a permissive syscall profile and unlimited locked memory for
development. Production deployments can use a custom seccomp profile and a
finite locked-memory budget that permit the server's required operations. See
[Linux tuning](/docs/server/linux-tuning#runtime-access-and-process-limits).
```yaml
services:
@@ -64,9 +64,9 @@ docker run -d --name iggy \
### Why these capabilities?
-- **`SYS_NICE`** - required for setting CPU affinity (`sched_setaffinity`) in
the thread-per-core architecture
-- **`seccomp:unconfined`** - required for `io_uring` syscalls which are
blocked by Docker's default seccomp profile
-- **`memlock: -1`** - `io_uring` needs to lock memory pages shared between
user space and kernel
+- **`SYS_NICE`** broadens scheduling and NUMA permissions. Pinning a thread
owned by the process with `sched_setaffinity` does not by itself require this
capability; container policies can impose additional restrictions. See [Linux
affinity
permissions](https://man7.org/linux/man-pages/man2/sched_setaffinity.2.html).
+- **`seccomp:unconfined`** permits `io_uring` calls blocked by Docker's
default seccomp profile. A custom profile allowing the required syscalls is an
alternative to disabling filtering. See [Docker seccomp
profiles](https://docs.docker.com/engine/security/seccomp/).
+- **`memlock: -1`** removes the process's locked-memory limit. A finite limit
is valid if it covers the runtime's requirements; neither setting overrides the
container's memory limit.
### Available images
diff --git a/content/docs/server/durability.mdx
b/content/docs/server/durability.mdx
new file mode 100644
index 00000000..7eab1254
--- /dev/null
+++ b/content/docs/server/durability.mdx
@@ -0,0 +1,150 @@
+---
+title: Durability
+description: "What a completed write guarantees, how to select topic
durability, and what survives a single-node crash."
+---
+
+A successful write has the guarantee selected by the topic's `durability`
+option. The default, `replicated`, waits for consensus commit and local
+application without an additional stable-storage barrier. `persisted` also
+requires a recoverable copy on stable storage at the replication quorum.
+
+**Both policies write data to disk.** The difference is what must have
+completed before the server reports success. On a single node the quorum is
+one, so `replicated` provides no second copy and can acknowledge messages
+that exist only in the server process's memory. For multiple replicas, see
+[Cluster Durability](/docs/clustering/durability).
+
+## Select the policy when creating the topic
+
+| Topic option | Default | Controls |
+| --- | --- | --- |
+| `durability` | `replicated` | Message production |
+| `consumer_offset_durability` | `replicated` | Explicit consumer-offset
stores and deletes |
+
+Both accept `replicated` or `persisted`. They default **independently**:
+setting `durability=persisted` leaves consumer offsets at `replicated`
+unless that option is also supplied. Neither inherits the other.
+
+For persisted messages and consumer offsets, create a topic in an existing
+stream:
+
+```bash
+iggy topic create my-stream my-topic 1 none \
+ --durability persisted \
+ --consumer-offset-durability persisted
+```
+
+These are [topic creation options](/docs/server/topic-options), not
+`config.toml` settings or server environment overrides. `GetTopic` reports
+the effective values. Both policies are create-only; `UpdateTopic` cannot
+change them on an existing topic.
+
+The former `enforce_fsync` topic option is rejected. Use `durability` for
+messages and `consumer_offset_durability` for offsets. The old server-wide
+`consumer_offset_enforce_fsync` setting has also been removed.
+
+## The single-node write path
+
+With `replicated`, the server commits and applies the operation locally.
+Message batches can remain in the in-memory partition journal after success.
+Ordinary segment writes still happen when a flush trigger fires, and a write
+that triggers a flush can wait for that I/O, but success does not require
+stable-storage synchronization.
+
+With `persisted`, successful message completion follows this path:
+
+```text
+local commit
+ -> write committed batches to .log and .index
+ -> synchronize both files and any newly created segment names
+ -> reply to the producer
+```
+
+The committed batches are flushed even below the ordinary flush thresholds.
+**Setting `durability=persisted` is sufficient; setting
+`messages_required_to_save=1` is not required.** A single-replica partition
+does not use the cluster's prepare WAL.
+
+For explicit offset changes, `consumer_offset_durability=persisted`
+similarly waits for the required local file and directory synchronization
+before success. On a singleton, persisting an offset does not force earlier
+`replicated` messages to be flushed. Select both policies as `persisted`
+when both messages and their stored offsets need that protection.
+
+## Flush scheduling is separate
+
+| Topic option | Default | Effect |
+| --- | --- | --- |
+| `messages_required_to_save` | `1024` | Message-count trigger for ordinary
segment writes |
+| `size_of_messages_required_to_save` | `1 MiB` | Byte-count trigger for
ordinary segment writes |
+| `segment_size` | `1 GiB` | Soft segment limit; a whole batch can cross it |
+
+A count or byte trigger, or a full active segment, can flush the committed
+journal prefix. Required persistence, capacity pressure, and lifecycle work
+can flush earlier. These thresholds are not a maximum data-loss guarantee
+and do not specify a periodic flush interval. A quiet `replicated` topic
+can retain messages in process memory while waiting for a trigger.
+
+Reducing the thresholds changes batching and I/O frequency. It does not give
+`replicated` the stable-storage guarantee of `persisted`.
+
+## What survives a failure
+
+This table describes acknowledged messages on one node, with no later
+retention, purge, or deletion:
+
+| Failure | `replicated` | `persisted` |
+| --- | --- | --- |
+| Process crash, `SIGKILL`, panic, or OOM kill | Unflushed messages can be
lost | Acknowledged messages are recoverable from local storage |
+| Machine crash, kernel panic, or power loss | Unflushed messages and
unsynchronized file writes can be lost | Acknowledged messages are recoverable
if the storage honors synchronization |
+| Storage device lost or irreparably damaged | No surviving local copy is
guaranteed | No surviving local copy is guaranteed |
+
+Completed buffered file writes normally remain in Linux's page cache after
+the server process dies. Messages still in the process's journal do not.
+Page cache alone does not protect against machine or power failure.
+
+A graceful shutdown drains work and forces a final flush of committed
+messages, regardless of the thresholds. This depends on the shutdown
+finishing and its I/O succeeding. A container or service termination that
+expires its grace period and sends `SIGKILL` is a process crash, not a
+completed graceful shutdown.
+
+## Which responses prove completion
+
+The policy applies to the completed server operation, not merely to enqueueing
+a request in a producer or receiving an early dispatch response.
+
+For HTTP production:
+
+| Request | Response | Meaning |
+| --- | --- | --- |
+| Default, or `?ack=replicated` | `201 Created`, `Iggy-Durability: replicated`
or `persisted` | The write completed under the topic's message policy |
+| `?ack=none` | `202 Accepted`, `Iggy-Durability: none` | Dispatch was
accepted; commit and persistence were not awaited |
+
+The query value `ack=replicated` selects the awaited path even when the
+topic uses `persisted`. It does not override the topic policy.
+
+**Poll auto-commit is asynchronous.** A successful poll does not confirm
+that its offset write committed or became durable, even with
+`consumer_offset_durability=persisted`. When processing needs an explicit
+offset completion guarantee, store the offset explicitly and await success.
+
+## Recovery and cost
+
+Recovery validates segment batches and partition identity. It can truncate
+an incomplete tail and rebuild a missing or torn sparse index. Interior
+damage, gaps, or intact records beyond damaged data can require refusal or
+peer repair instead of truncation. A single node cannot fetch a missing copy
+from another replica. See [Storage Engine](/docs/server/storage-engine).
+
+`persisted` adds synchronization to the completion path. Its latency and
+throughput depend on storage, filesystem, batching, and concurrency. Measure
+both policies on the intended hardware rather than applying results from the
+removed `enforce_fsync` configuration. The
+[benchmarking guide](/docs/server/benchmarking) explains how to reproduce a
+workload.
+
+The implementation references are the
+[topic policy
definitions](https://github.com/apache/iggy/blob/97f7b0c0335f81691bef34f923f21d783016a6a4/core/common/src/types/options/durability.rs),
+[partition completion
paths](https://github.com/apache/iggy/blob/97f7b0c0335f81691bef34f923f21d783016a6a4/core/partitions/src/iggy_partition.rs),
+and [HTTP response
handling](https://github.com/apache/iggy/blob/97f7b0c0335f81691bef34f923f21d783016a6a4/core/server/src/http/handlers.rs).
diff --git a/content/docs/server/linux-tuning.mdx
b/content/docs/server/linux-tuning.mdx
new file mode 100644
index 00000000..12395e66
--- /dev/null
+++ b/content/docs/server/linux-tuning.mdx
@@ -0,0 +1,145 @@
+---
+title: Linux tuning
+description: "Prepare Linux hosts for Iggy and measure the effects of memory,
CPU, storage, and network tuning."
+---
+
+Host settings affect both production latency and benchmark results. Start with
the distribution defaults, satisfy Iggy's runtime requirements, then change one
setting at a time. The examples below are experiments for dedicated Linux
hosts, not a configuration to apply unchanged to every VM.
+
+Use the [benchmarking guide](/docs/server/benchmarking) to compare throughput,
p50, p99 and p99.9 latency, memory use, and errors under the same workload.
Keep the server version, transport, topic durability, CPU allocation, and
storage identical between runs.
+
+## Inspect the host first
+
+Run these read-only checks before changing the VM image or starting a
benchmark:
+
+```bash
+uname -r
+lscpu
+free -h
+swapon --show
+ulimit -Sn
+ulimit -Hn
+ulimit -Sl
+ulimit -Hl
+sysctl vm.swappiness vm.dirty_background_bytes vm.dirty_bytes
+sysctl vm.dirty_background_ratio vm.dirty_ratio
+grep -E '^(HugePages_|Hugepagesize|Hugetlb|AnonHugePages)' /proc/meminfo
+ls /sys/kernel/mm/hugepages/
+```
+
+The shell's limits are not necessarily the service's limits. For a running
process, inspect `/proc/PID/limits`, `/proc/PID/status`, and its cgroup limits.
Replace `PID` with that process's ID. Record the VM type, vCPU topology, memory
limit, disk type and provisioned throughput, filesystem, and network bandwidth
with the result.
+
+## Runtime access and process limits
+
+Iggy uses `io_uring`. The kernel must implement the operations required by the
server build, and the service or container must be allowed to use them. The
server's startup diagnostics identify failed ring creation and unsupported
operations. Where the kernel exposes them, inspect:
+
+```bash
+sysctl kernel.io_uring_disabled kernel.io_uring_group
+```
+
+`io_uring_disabled=0` permits ring creation. With `1`, an unprivileged service
needs membership in the configured `io_uring_group`. With `2`, creation is
disabled. Container syscall filters can independently deny access. Configure
access for the service rather than assuming that a root shell proves the
container can start. See the [kernel's io_uring
controls](https://docs.kernel.org/admin-guide/sysctl/kernel.html#io-uring-disabled)
and [Docker deployment guidance](/docs/server/docker).
+
+Size file-descriptor limits for connections, partitions, open segment files,
and runtime overhead. `LimitNOFILE=65536` is an example starting budget, not an
Iggy requirement. Increase it if measurements justify more. A systemd service
can use the following drop-in:
+
+```ini
+[Service]
+LimitNOFILE=65536
+LimitMEMLOCK=infinity
+```
+
+Unlimited locked memory is an option for a dedicated service with a planned
memory budget. A finite limit is also valid if it covers the runtime's
locked-memory requirements. Neither setting creates RAM or overrides a
container's memory ceiling.
+
+Recheck the actual process limits after restarting the service. PAM settings
in `/etc/security/limits.conf` do not configure a system service. See [systemd
resource
limits](https://man7.org/linux/man-pages/man5/systemd.exec.5.html#PROCESS_PROPERTIES).
+
+## Swapping and writeback
+
+On a host with disk-backed swap and enough RAM for the workload, try
`vm.swappiness=10` and compare it with the existing value. It reduces the
preference for swapping relative to reclaiming file cache. It does not disable
swapping, and even `0` permits swapping under sufficient pressure. zram and
zswap can warrant different values. See [Linux
swappiness](https://docs.kernel.org/admin-guide/sysctl/vm.html#swappiness).
+
+```bash
+# Optional experiment. Record the previous value first.
+sudo sysctl -w vm.swappiness=10
+```
+
+Leave memory for the kernel, page cache, other processes, and Iggy's buffers.
Watch swap-in/out with `vmstat 1` and memory pressure in
`/proc/pressure/memory`. Disabling swap does not fix an undersized memory
budget.
+
+Dirty-page limits govern buffered writeback. Large percentage limits, such as
`dirty_background_ratio=10` and `dirty_ratio=30`, can permit large bursts on
high-memory VMs. If writeback coincides with latency spikes, compare explicit
byte limits sized from disk throughput and the acceptable backlog.
+
+Byte and ratio forms are alternatives. Writing one disables its corresponding
other form. These settings do not replace Iggy's persistence barriers. See
[Linux writeback
controls](https://docs.kernel.org/admin-guide/sysctl/vm.html#dirty-background-bytes).
+
+## Huge pages and mimalloc
+
+Huge pages can reduce translation overhead, but they consume memory
differently from ordinary pages. Measure both throughput and tail latency.
Iggy's default server build uses mimalloc. Check the startup allocator message,
since builds that disable mimalloc do not use its environment options.
+
+There are three distinct controls:
+
+| Control | Meaning |
+| --- | --- |
+| Transparent huge pages (THP) | Kernel promotion of eligible mappings.
Inspect the top-level and per-size THP policies. |
+| Linux HugeTLB pools | Explicitly reserved pages. `vm.nr_hugepages` counts
pages of the default size reported by `Hugepagesize`. |
+| mimalloc environment options | Tell the allocator which page mechanisms to
request. They are not kernel pool sizes. |
+
+### Explicit large pages
+
+First verify the default page size and leave sufficient ordinary memory. If
`Hugepagesize` is `2048 kB`, 512 reserved pages consume 1 GiB:
+
+```bash
+# Only for a host whose default Hugepagesize is 2048 kB
+# and whose memory budget permits a 1 GiB reservation.
+sudo sysctl -w vm.nr_hugepages=512
+grep -E '^(HugePages_|Hugepagesize|Hugetlb)' /proc/meminfo
+```
+
+The requested count may not be allocated on a fragmented host. Boot-time
reservation is more reliable. Reserved HugeTLB memory is unavailable for
ordinary allocations or page cache, even while unused. Check NUMA placement and
the actual pool, not just the requested count. See [HugeTLB
administration](https://docs.kernel.org/admin-guide/mm/hugetlbpage.html).
+
+Try the allocator option on a mimalloc-enabled server, with access to the pool
configured for its service account:
+
+```bash
+MIMALLOC_ALLOW_LARGE_OS_PAGES=1 MIMALLOC_VERBOSE=1 ./target/release/iggy-server
+```
+
+Set allocator variables in the launching environment or the service's
`Environment=` entries before process startup. The server can allocate before
loading its `.env` file. Verbose allocator output is useful for checking the
experiment, then turn it off for timed runs.
+
+**Do not copy a Linux page count into `MIMALLOC_RESERVE_HUGE_OS_PAGES`.** That
mimalloc option counts **1 GiB pages**. A value of `2048` requests 2 TiB,
whereas `vm.nr_hugepages=2048` represents 4 GiB only when the default HugeTLB
page size is 2 MiB. A 2 MiB pool does not provide a matching 1 GiB reservation.
+
+A separate 1 GiB-page experiment requires platform support, a suitable pool,
and a budget for each process. mimalloc advises normally choosing this
reservation mode separately from `MIMALLOC_ALLOW_LARGE_OS_PAGES`. See [mimalloc
environment
options](https://github.com/microsoft/mimalloc#environment-options). HugeTLB
allocation can require `CAP_IPC_LOCK` or membership in `vm.hugetlb_shm_group`,
as described in [mmap
permissions](https://man7.org/linux/man-pages/man2/mmap.2.html).
+
+### Transparent huge pages
+
+Inspect the current policy instead of universally enabling or disabling THP:
+
+```bash
+cat /sys/kernel/mm/transparent_hugepage/enabled
+cat /sys/kernel/mm/transparent_hugepage/defrag
+```
+
+Compare the existing policy with `madvise` if allocator page promotion or
compaction is relevant. Direct reclaim and compaction can stall allocations. On
recent kernels, per-size policies and explicit `MADV_COLLAPSE` requests also
matter, so the top-level switch alone does not describe all behavior. See
[transparent huge-page
controls](https://docs.kernel.org/admin-guide/mm/transhuge.html).
+
+Verify actual usage in `/proc/PID/smaps`: `AnonHugePages` reports PMD-sized
anonymous THP, while `Private_Hugetlb` and `Shared_Hugetlb` identify explicit
huge-page mappings. A successful pool reservation alone does not prove Iggy
used it.
+
+## CPU placement and NUMA
+
+Allocate dedicated CPUs where possible. On the current server, `[sharding]
cpu_allocation` and `pin_cores` control shard placement, and the process's
allowed CPU set constrains allocation. Check the embedded configuration for the
exact version being deployed. Separate server CPUs from a colocated load
generator and leave capacity for kernel and network work.
+
+Inspect NUMA topology with `lscpu` or `numactl --hardware`. Compare
physical-core-only placement with SMT when relevant. Match memory placement and
any per-node huge-page reservation to the CPUs running the shards. Avoid
assuming that allocating a pool on one NUMA node helps shards running on
another.
+
+For a reproducible dedicated-host experiment, compare the `performance` CPU
governor with the current governor if the guest exposes CPU frequency control.
Power and thermal limits can still constrain frequency. Many cloud guests
cannot control the host's governor. Record the setting rather than assuming a
command changed it. See [CPU performance
scaling](https://docs.kernel.org/admin-guide/pm/cpufreq.html).
+
+Inspect CPU steal time, cgroup throttling, and per-core utilization before
increasing shard counts. If one core handles most interrupts or receive
processing, inspect RSS queue distribution and IRQ affinity. Do not disable
irqbalance or assign every IRQ to a fixed core without measuring. See [Linux
network scaling](https://docs.kernel.org/networking/scaling.html).
+
+## Network and storage
+
+Socket-buffer requirements depend on bandwidth and round-trip time. For
example, 10 Gbit/s across 10 ms has a bandwidth-delay product of about 12.5 MB.
If the connection is window-limited, inspect `net.core.rmem_max`,
`net.core.wmem_max`, `net.ipv4.tcp_rmem`, and `net.ipv4.tcp_wmem`. Larger
ceilings permit larger buffers, but do not force the application to use them or
guarantee higher throughput.
+
+Raise `somaxconn`, `tcp_max_syn_backlog`, or `netdev_max_backlog` only when
backlog pressure or drops show the need. QUIC uses UDP, so TCP tuning does not
configure its flow-control windows or UDP receive buffers.
+
+Keep changes such as `tcp_tw_reuse`, shorter TCP timeouts, and wider
ephemeral-port ranges tied to a demonstrated connection-churn problem. See [TCP
controls](https://docs.kernel.org/networking/ip-sysctl.html) and [core
networking controls](https://docs.kernel.org/admin-guide/sysctl/net.html).
+
+Use storage with measured sustained bandwidth and latency, and record cloud
volume and instance limits. Keep free space for segments, WAL, and recovery.
Compare warm-cache and cold-cache runs explicitly. Do not periodically drop
caches or disable storage flushes to improve a benchmark number. Storage,
retention, and the selected durability policy must match the workload being
evaluated.
+
+## Apply, verify, and reproduce
+
+Trial a change with a runtime setting first. Keep the previous value for
rollback. Persist accepted sysctl values in a dedicated `/etc/sysctl.d/` file
and accepted service limits or allocator variables in that service's systemd
drop-in. THP sysfs writes require a boot-time mechanism of their own. Recheck
effective values after reboot and process restart.
+
+Collect `vmstat 1`, `iostat -xz 1`, `mpstat -P ALL 1`, pressure counters, and
network drop/retransmission counters alongside benchmark results. The iostat
and mpstat tools are provided by the sysstat package. Run long enough to
include segment flushes, WAL checkpoints, and sustained device limits, rather
than measuring only buffered admission.
+
+Iggy's topic `durability` and `consumer_offset_durability` independently
default to `replicated`. `persisted` changes completion guarantees, and huge
pages or sysctl settings do not replace it. Poll auto-commit remains
asynchronous. Record both policies, replication-group size, payload and batch
sizes, warmup, CPU allocation, and the exact host changes with each benchmark
result.
+
diff --git a/content/docs/server/meta.json b/content/docs/server/meta.json
index 4c9c3bd6..a80c2e9d 100644
--- a/content/docs/server/meta.json
+++ b/content/docs/server/meta.json
@@ -1,4 +1,4 @@
{
"title": "Server",
- "pages": ["introduction", "configuration", "topic-options",
"storage-engine", "networking", "security", "docker", "benchmarking"]
+ "pages": ["introduction", "configuration", "topic-options",
"storage-engine", "durability", "networking", "security", "docker",
"linux-tuning", "benchmarking"]
}
diff --git a/content/docs/server/security.mdx b/content/docs/server/security.mdx
index 98fe47f3..fa638600 100644
--- a/content/docs/server/security.mdx
+++ b/content/docs/server/security.mdx
@@ -107,10 +107,10 @@ For production deployments, provide proper certificates
via `cert_file` and `key
## Data encryption at rest
-Iggy supports optional **AES-256-GCM** encryption for message payloads and
state commands. When enabled, all data is encrypted before being written to
disk and decrypted when read. The encryption key must be a 32-byte
base64-encoded string.
+Iggy supports optional **AES-256-GCM** encryption for message payloads and
user headers. They are encrypted before storage and decrypted for polling.
Metadata journals, metadata snapshots, and structural record headers remain
unencrypted. This option does not provide whole-directory encryption. The
encryption key must decode from base64 to 32 bytes.
```toml
-[system.encryption]
+[encryption]
enabled = false
key = "" # 32-byte base64-encoded key
```
diff --git a/content/docs/server/storage-engine.mdx
b/content/docs/server/storage-engine.mdx
index 3deee01b..a9fcca05 100644
--- a/content/docs/server/storage-engine.mdx
+++ b/content/docs/server/storage-engine.mdx
@@ -11,7 +11,7 @@ Iggy's storage engine is built around the concept of a
**segmented append-only l
## Directory layout
-All data lives under the `system.path` directory (default `local_data`):
+All data lives under the root `path` directory (default `local_data`,
overridden with `IGGY_PATH`):
```bash
local_data/
@@ -32,6 +32,7 @@ local_data/
└── {partition_id}/
├── superblock.a # Partition
consensus superblock pair
├── superblock.b
+ ├── prepares-{created_revision}/ # Prepare WAL, when
a cluster topic policy is persisted
├── offsets/
│ ├── consumers/ # Stored offsets of
individual consumers
│ └── groups/ # Stored offsets of
consumer groups
@@ -41,7 +42,7 @@ local_data/
└── 00000000000016000000.index
```
-Stream, topic, and partition ids are numeric and **0-based**. Each partition
directory holds pairs of `.log` and `.index` files. The filename is the start
offset of the segment's first message, zero-padded to 20 digits. Next to the
segments, every partition keeps its own superblock pair (replica identity and
consensus state for that partition) and an `offsets/` tree for consumer offset
storage.
+Stream, topic, and partition ids are numeric and **0-based**. Each partition
directory holds pairs of `.log` and `.index` files. The filename is the start
offset of the segment's first message, zero-padded to 20 digits. Next to the
segments, every partition keeps its own superblock pair (replica identity and
consensus state for that partition) and an `offsets/` tree for consumer offset
storage. Multi-replica partitions also keep a prepare WAL when either topic
durability policy is `persisted`.
## Segmented log
@@ -80,56 +81,51 @@ There is no index caching configuration: the in-memory
index cache is an interna
## Write pipeline
-Messages are admitted, stamped, and buffered in the partition journal, then
replicated; committed batches reach a segment once a flush trigger fires:
+Messages are admitted, stamped, and buffered in the partition journal. Their
completion path depends on the topic's durability policy:
```mermaid
graph TD
A["Client SendMessages: batch with producer-computed checksums"] -->
B["Admission: verify checksums, stamp partition_id"]
B --> C["Stamp base_offset / base_timestamp, recompute batch_checksum"]
C --> D["Append to the in-memory partition journal"]
- D --> E["Partition consensus: prepare replicated to the partition's
replicas"]
- E --> F{"Quorum reached?"}
- F -->|"Yes"| G["Commit"]
- G --> H{"Flush trigger reached?"}
- H -->|"No"| W["Stay buffered until a later commit trips the gate"]
- H -->|"Yes"| I["MessagesWriter: vectored write to .log + sparse index entry"]
- I --> J{"enforce_fsync?"}
- J -->|"Yes"| K["fdatasync .log + .index"]
- J -->|"No"| L["OS writeback"]
- K --> M{"Segment past segment_size?"}
- L --> M
- M -->|"Yes"| N["Seal segment, open next"]
- M -->|"No"| O["Done"]
+ D --> E{"Multiple replicas?"}
+ E -->|"Yes"| R["Replicate prepares; meet required WAL barriers before
PrepareOk"]
+ E -->|"No"| F["VSR quorum commit"]
+ R --> F
+ F --> G["Apply committed state; write segments when required"]
+ G --> H{"Single replica and persisted?"}
+ H -->|"Yes"| J["Complete local segment synchronization"]
+ H -->|"No"| I["Reply under the topic's completion policy"]
+ J --> I
```
-The flush thresholds and fsync policy are **per-topic creation options** (they
used to be server-wide config):
+The flush thresholds and completion policies are **per-topic creation
options**:
- `messages_required_to_save` (default `1024`): count threshold.
- `size_of_messages_required_to_save` (default `1 MiB`): byte threshold.
-- `enforce_fsync` (default `false`): fdatasync the `.log` and `.index` on
every flush.
+- `durability` (default `replicated`): message completion policy; `persisted`
requires recoverable stable storage at the replication quorum.
+- `consumer_offset_durability` (default `replicated`): the independent policy
for explicit offset stores and deletes.
-A third trigger is independent of these options: when buffered journal bytes
reach the active segment's `segment_size`, the flush fires. Whichever trigger
trips first flushes. These are soft limits: a flush writes whole batches, so
the actual flushed count or size can overshoot. The gate is evaluated on commit
events, so a sub-threshold batch stays buffered until a later commit trips it
(shutdown flushes unconditionally). The `MessagesWriter` uses **vectored I/O**
(`writev`) with up to ` [...]
+A full active segment also triggers an ordinary flush. These are soft limits:
a flush writes whole committed batches, so the actual count or size can
overshoot. Required persistence, capacity pressure, and lifecycle operations
can flush below the thresholds. The settings do not define a periodic flush
interval. The `MessagesWriter` uses **vectored I/O** with up to 1024 buffers
per syscall, so many buffered batches can land in one write.
-Every partition belongs to a consensus group (Viewstamped Replication), and
writes are prepared and quorum-acknowledged before commit. Stamping happens at
journal-append time, before replication acks: the primary stamps `base_offset`
from the log position and `base_timestamp` from the prepare timestamp, forwards
the stamped bytes verbatim, and each backup re-derives the expected pair and
refuses a mismatch, so the batch (and its recomputed `batch_checksum`) is
byte-identical across repli [...]
+With `replicated`, completion does not wait for an additional stable-storage
barrier. With `persisted`, a single-replica partition flushes and synchronizes
committed segment state before success. A multi-replica partition instead gates
the required prepare acknowledgements on recoverable WAL history, so persisted
success can precede an ordinary segment flush. `messages_required_to_save=1` is
not required.
+
+For multi-replica partitions, the disk prepare WAL is enabled when either
policy is `persisted` and includes message payloads even if only the offset
policy is persisted. Checkpoints synchronize materialized files before
reclaiming WAL history. See [Durability](/docs/server/durability) and [Cluster
Durability](/docs/clustering/durability) for completion guarantees and failure
behavior.
## Boot-time segment recovery
-At startup the server recovers every partition from its files, tolerating a
torn tail from a crash:
+At startup, recovery validates batch checksums, partition identity, and the
segment chain before serving data. With message durability `replicated`, it
walks the log from byte zero because buffered writeback can preserve later
pages before earlier ones. With `persisted`, completed durable segment flushes
allow an index-anchored recovery path.
-1. Sweep leftovers: `.staging` files and orphan `.index` files (an index
without its `.log`) are deleted. A `.log` is **never deleted**.
-2. For each segment, read the segment's bounds (offsets, timestamps, size)
from the 24-byte sparse index.
-3. If the index is missing or torn, walk the `.log` batch by batch and recover
the bounds from the records themselves. This isn't tail-only: with the default
`enforce_fsync = false`, a torn index can appear mid-chain, and the log walk
recovers it.
-4. If neither the index nor the log holds one whole batch, the segment is
recovered as empty: sizes are zeroed so the next append overwrites the torn
bytes instead of stranding undecodable garbage inside the readable range.
-5. The last segment is left unsealed and becomes the active segment.
+A missing or torn sparse index can be rebuilt from valid log data. An
incomplete tail can be truncated, but an interior gap, intact records after
damage, or a contradiction with durable history can require recovery refusal.
Recovery does not blindly truncate at the first invalid batch.
-An index entry pointing past the end of its `.log` file (a torn write
mid-chain that cannot be reconciled) **refuses recovery** for that partition
rather than serving corrupt data.
+When a prepare WAL exists, recovery reconciles it with materialized segment
and offset state. Unrecoverable partition data is fenced and repaired from
peers when available. It must not be replaced by an empty healthy partition. A
singleton has no peer copy to fetch.
## Read integrity
Disk reads are verified before they reach a consumer:
```toml
-[system.partition]
+[partition]
validate_checksum = true
```
@@ -142,7 +138,7 @@ Polls serve stored batch records as-is (a reply may be a
server-sliced view of a
Iggy includes a custom memory pool to eliminate allocation overhead on the hot
path. The pool has **28 buckets** with buffer sizes from 4 KiB up to 512 MiB
(non-uniform spacing, denser around common message sizes, with sizes above 2
MiB rounded to hugepage-friendly steps). Components request a buffer from the
appropriate bucket and return it when done.
```toml
-[system.memory_pool]
+[memory_pool]
enabled = true
size = "4 GiB" # Total pool size (minimum 512 MiB, multiple of the
4096-byte page size)
bucket_capacity = 8192 # Buffers per bucket (power of 2, minimum 128)
@@ -178,25 +174,26 @@ Partition directories carry the same superblock mechanism
for their own consensu
## Encryption
-Iggy supports optional **AES-256-GCM** encryption for message payloads and
state commands. When enabled, data is encrypted before being written to disk
and decrypted when read. The key is a 32-byte, base64-encoded string.
+Iggy supports optional **AES-256-GCM** encryption for message payloads and
user headers. They are encrypted before storage and decrypted for polling.
Metadata journals, metadata snapshots, and structural record headers remain
unencrypted. The key must decode from base64 to 32 bytes. See
[Security](/docs/server/security#data-encryption-at-rest).
```toml
-[system.encryption]
+[encryption]
enabled = false
key = "" # 32-byte base64-encoded key
```
## Compression
-The `compression_algorithm` topic option accepts `none` (default) and `gzip`,
but it is a **placeholder today**: the value is persisted and reported back,
and no compression is applied anywhere - segments store payloads exactly as
sent. To compress today, do it client-side and tag messages via user headers -
see the [message headers
examples](https://github.com/apache/iggy/tree/master/examples/rust/src/message-headers)
in the Iggy repo.
+The `compression_algorithm` topic option accepts `none` (default) and `gzip`,
but it is a **placeholder today**: the value is persisted and reported back,
but no message compression is applied. Payload encryption still applies when
enabled. To compress today, do it client-side and tag messages via user headers
- see the [message headers
examples](https://github.com/apache/iggy/tree/master/examples/rust/src/message-headers)
in the Iggy repo.
## Removed and relocated settings
Earlier releases documented several storage features and `[system.*]` keys
that no longer exist:
- **`cache_indexes`**: removed. Index caching is internal now (see
[Indexes](#indexes)).
-- **`[system.message_deduplication]`**: the server-side deduplicator was
removed entirely.
-- **Segment archiving / S3 backup**: removed. `[system.segment]
archive_expired` must stay `false`. Setting it to `true` aborts boot.
-- **`[system.topic]` / `[system.partition]` / `[system.segment]` storage
knobs** (`max_size`, `message_expiry`, `enforce_fsync`,
`messages_required_to_save`, `size_of_messages_required_to_save`, `size`,
`preallocate`): relocated to per-topic options.
+- **`[system.message_deduplication]`**: removed configuration; it does not
configure the partition request deduplication table.
+- **Segment archiving / S3 backup placeholders**: removed. Delete
`archive_expired` and the old `[system.segment]` table rather than retaining a
`false` value.
+- **`enforce_fsync`**: replaced by the topic `durability` policy.
Consumer-offset completion has its own `consumer_offset_durability` policy.
Both default independently to `replicated`.
+- **`[system.*]` tables**: removed or moved to root tables. Retention, segment
size, preallocation, and flush thresholds are per-topic options.
-The config loader **refuses to boot** while any relocated or removed key is
still set, in the file or the environment, so stale configs fail fast instead
of silently ignoring a knob. The full mapping is in [Relocated configuration
keys](/docs/server/configuration#relocated-configuration-keys).
+The config loader rejects the old `[system]` table and `IGGY_SYSTEM_*`
environment mappings. The full mapping is in [Relocated configuration
keys](/docs/server/configuration#relocated-configuration-keys).
diff --git a/content/docs/server/topic-options.mdx
b/content/docs/server/topic-options.mdx
index c75df021..e32c42b8 100644
--- a/content/docs/server/topic-options.mdx
+++ b/content/docs/server/topic-options.mdx
@@ -15,13 +15,16 @@ Options are key-value pairs sent with `CreateTopic`.
Unknown keys are **rejected
| `message_expiry` | none | | Delete sealed segments older than this. |
| `compression_algorithm` | `none` | `none` or `gzip` | Placeholder: stored
and reported, no compression applied yet. |
| `segment_size` | 1 GiB | 512-byte multiple, at least 1 MiB, at most 1 GiB |
Soft size limit per segment: a segment may close one whole batch past it. |
-| `enforce_fsync` | `false` | | fsync every write to this topic's partitions. |
+| `durability` | `replicated` | `replicated` or `persisted` | Message
completion policy. `persisted` requires recoverable stable-storage copies at
the replication quorum before success. |
+| `consumer_offset_durability` | `replicated` | `replicated` or `persisted` |
Completion policy for explicit consumer-offset stores and deletes, independent
of message durability. |
| `messages_required_to_save` | 1024 | non-zero, at most 16777216 | Flush the
journal once it holds this many messages. |
| `size_of_messages_required_to_save` | 1 MiB | at most 1 GiB | Flush the
journal once it holds this many bytes. Paired with the message count; whichever
threshold trips first flushes. |
| `preallocate_segments` | `false` | `segment_size` x partitions at most 64
GiB per create | Reserve each segment's bytes up front where the filesystem
supports it. |
Both retention policies can be active at once. The active segment is **never
touched**. Deletion is done by the server's segment cleaner
(`[data_maintenance.messages]`, enabled by default).
+Both durability policies write data to disk and default independently to
`replicated`. Neither inherits the other. Flush thresholds schedule ordinary
segment writes; required persistence, capacity pressure, or lifecycle
operations can flush earlier. They do not weaken the `persisted` completion
guarantee.
+
Value forms are forgiving: byte sizes accept a raw number of bytes or a string
like `"128 MiB"`, expiry accepts microseconds or a humantime string like `"7
days"`, booleans accept `true`/`false`. Create admission re-parses and
re-encodes what you send, so a string `segment_size=128MiB` is stored as the
number it names.
`preallocate_segments` reserves exactly `segment_size` of real disk per
partition the moment the topic is created (and again as segments rotate). With
the default 1 GiB segment size that's 1 GiB per partition up front, which is
why it's opt-in and why one create is **capped at 64 GiB** of total reservation.
@@ -35,14 +38,15 @@ At creation, every interface takes the same keys:
# partitions count, compression algorithm.
iggy topic create my-stream my-topic 1 none \
--set segment_size=128MiB \
- --set enforce_fsync=true
+ --durability persisted \
+ --consumer-offset-durability persisted
```
-In the CLI, `compression_algorithm`, `message_expiry`, and `max_topic_size`
also have first-class parameters on `iggy topic create`. `--set` covers the
storage knobs that have no named parameter of their own. Values are sent as
strings and parsed server-side, so `iggy options topic` (below) tells you
exactly what this server accepts.
+In the CLI, both durability policies have named flags. `--set` is repeatable
and covers other storage options, such as `segment_size`.
`compression_algorithm`, `message_expiry`, and `max_topic_size` also have
first-class parameters. Run `iggy options topic` to discover what the server
accepts.
-SDKs pass options as a parameter on the create call (`TopicCreateOptions::raw`
in Rust, an `options` dictionary or list in Node, Python, Go, Java, C#, and
C++). The HTTP API takes them as a plain string map in the create body.
+SDKs expose typed durability values in their topic creation options. In Rust,
set `TopicCreateOptions::durability` and
`TopicCreateOptions::consumer_offset_durability` to `Durability::Replicated` or
`Durability::Persisted`. The HTTP API takes `"durability"` and
`"consumer_offset_durability"` as string values in the create body's `options`
map.
-Keys you don't send are resolved to the defaults above by the admitting server
and stored as **derived** entries. `GetTopic` returns both blocks, explicit and
derived, so the effective value of every knob is always visible, along with who
chose it.
+Keys absent from the wire request are resolved by the admitting server and
stored as **derived** entries. Typed SDKs can send their default durability
values explicitly. `GetTopic` returns explicit and derived blocks, so the
effective values and the provenance of the request remain visible.
## Create-only vs updatable
@@ -54,7 +58,7 @@ Keys you don't send are resolved to the defaults above by the
admitting server a
Updates are **patches**: a key you don't send keeps its current value.
-The storage knobs (`segment_size`, `enforce_fsync`,
`messages_required_to_save`, `size_of_messages_required_to_save`,
`preallocate_segments`) are **create-only**. They describe how a partition's
storage is laid down: changing `segment_size` mid-segment would leave segments
sized by different caps, and preallocation can only act on a file not yet
opened. A topic gets them at creation and keeps them, so its segments stay
uniform. This is also why the server refuses to boot on the old confi [...]
+The storage options (`segment_size`, `durability`,
`consumer_offset_durability`, `messages_required_to_save`,
`size_of_messages_required_to_save`, `preallocate_segments`) are
**create-only**. A topic gets them at creation and keeps them. `UpdateTopic`
rejects changes to either durability policy, so select both before creating the
topic.
## Discovering the catalog
@@ -74,6 +78,8 @@ Discovery matters most on the binary transports: TCP, QUIC,
and WebSocket carry
## Durability
-Single-node durability is a per-topic decision. For the strongest guarantee,
create the topic with `enforce_fsync=true`. Every write is then **fsynced
before it is acknowledged**. Without it, the flush thresholds
(`messages_required_to_save`, `size_of_messages_required_to_save`) bound how
much buffered data a crash can lose, at much higher throughput.
+`replicated` waits for quorum commit and application without an additional
stable-storage barrier. `persisted` also requires recoverable stable-storage
copies at the required quorum. On a single node that means recoverable local
persistence before success; in a cluster it uses durable prepare history.
Setting `messages_required_to_save=1` is not required for `persisted`.
+
+The former `enforce_fsync` option is rejected. Use `durability` for messages
and `consumer_offset_durability` for explicit offset changes. Poll auto-commit
remains asynchronous, and HTTP `ack=none` confirms dispatch only, regardless of
these policies.
-In cluster mode, durability comes from quorum replication: a write is
acknowledged once a majority of replicas hold it. See
[Clustering](/docs/clustering/vsr).
+See [Durability](/docs/server/durability) for single-node failure behavior and
[Cluster Durability](/docs/clustering/durability) for the quorum and WAL
guarantees.