numinnex opened a new pull request, #3976:
URL: https://github.com/apache/iggy/pull/3976
`FileStorage::truncate` is the boot-time repair for a torn metadata WAL tail.
It used compio's `set_len`, which submits `IORING_OP_FTRUNCATE`. That opcode
landed in kernel 6.9. Below it the driver probes the opcode as unsupported
and
falls back to `push_blocking`, but shard proactors are built with
`thread_pool_limit(0)`, so the fallback panics with `the thread pool is
needed
but no worker thread is running`. The panic fires inside dispatch, outside
`catch_unwind_io`.
Net effect: a crash that tore the last WAL append made the next boot kill the
shard instead of repairing it, and the repair is idempotent, so the node
stayed
down across restarts.
Two conditions have to coincide, so this is not a routine path:
```text
kernel < 6.9 IORING_OP_FTRUNCATE absent, compio falls back
AND
torn WAL tail crash mid-append, so boot calls truncate_or_fail
The affected range is 5.19 through 6.8, not everything below 6.9. Ring setup
already requires IORING_SETUP_COOP_TASKRUN and IORING_SETUP_TASKRUN_FLAG,
which need 5.19, so RHEL 9 (5.14) and stock Ubuntu 22.04 (5.15) never start
the
server at all and were never exposed. What this actually broke is Debian 12
and
AL2023 (6.1), and Ubuntu 24.04 and 22.04-HWE (6.8). macOS aarch64 is exempt
because create_shard_executor keeps a blocking pool there by design.
The fix
Truncate synchronously through std::fs on the stored path, which needs
neither the opcode nor the blocking pool. This mirrors what segment recovery
already does in truncate_to. The sync_all moves inside truncate, so the
repair is durable on its own and the caller no longer pairs it with a
separate
fsync. sync_all rather than sync_data because the file length is metadata,
and without it a power cut right after the repair re-presents the torn tail.
truncate is no longer async. An async fn that never awaits trips
clippy::unused_async, and the journal crate denies clippy::pedantic.
Blocking the shard thread costs nothing here: the sole caller is boot-time
repair, before the shard serves traffic.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]