numinnex opened a new pull request, #3976:
URL: https://github.com/apache/iggy/pull/3976

   `FileStorage::truncate` is the boot-time repair for a torn metadata WAL tail.
   It used compio's `set_len`, which submits `IORING_OP_FTRUNCATE`. That opcode
   landed in kernel 6.9. Below it the driver probes the opcode as unsupported 
and
   falls back to `push_blocking`, but shard proactors are built with
   `thread_pool_limit(0)`, so the fallback panics with `the thread pool is 
needed
   but no worker thread is running`. The panic fires inside dispatch, outside
   `catch_unwind_io`.
   
   Net effect: a crash that tore the last WAL append made the next boot kill the
   shard instead of repairing it, and the repair is idempotent, so the node 
stayed
   down across restarts.
   
   Two conditions have to coincide, so this is not a routine path:
   
   ```text
   kernel < 6.9          IORING_OP_FTRUNCATE absent, compio falls back
     AND
   torn WAL tail         crash mid-append, so boot calls truncate_or_fail
   
   The affected range is 5.19 through 6.8, not everything below 6.9. Ring setup
   already requires IORING_SETUP_COOP_TASKRUN and IORING_SETUP_TASKRUN_FLAG,
   which need 5.19, so RHEL 9 (5.14) and stock Ubuntu 22.04 (5.15) never start 
the
   server at all and were never exposed. What this actually broke is Debian 12 
and
   AL2023 (6.1), and Ubuntu 24.04 and 22.04-HWE (6.8). macOS aarch64 is exempt
   because create_shard_executor keeps a blocking pool there by design.
   
   The fix
   
   Truncate synchronously through std::fs on the stored path, which needs
   neither the opcode nor the blocking pool. This mirrors what segment recovery
   already does in truncate_to. The sync_all moves inside truncate, so the
   repair is durable on its own and the caller no longer pairs it with a 
separate
   fsync. sync_all rather than sync_data because the file length is metadata,
   and without it a power cut right after the repair re-presents the torn tail.
   
   truncate is no longer async. An async fn that never awaits trips
   clippy::unused_async, and the journal crate denies clippy::pedantic.
   Blocking the shard thread costs nothing here: the sole caller is boot-time
   repair, before the shard serves traffic.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to