Really nice idea Chris!

One thing I think it’s worth flagging for the proposal:

this CEP introduces a new SSTable major version. That's a relevant change for 
anything outside nodetool/the core read path that parses SSTables directly. 
e.g. analytics library uses the concept of per major version bridge to 
deserialize.

With this approach, current bridges don't have a notion of "data doesn't start 
at logical offset zero” - they assume a chunk's first partition begins the 
SSTable's data - which is fine (this is a new version!). It'd help if the CEP 
explicitly calls out that third-party/off-node SSTable readers are a 
compatibility surface here, not just in-process Cassandra binaries.

Bernardo

From: Chris Lohfink <[email protected]>
Date: Monday, 14 September 2026 at 21:23
To: [email protected] <[email protected]>
Subject: [DISCUSS] CEP-66: Zero-copy SSTable splitting

Hi everyone,

I'd like to open CEP-66, Zero-copy SSTable splitting, for discussion:

https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting

Anticompaction and partial-range streaming currently rewrite rows whose encoded 
representation already exists on disk. This consumes CPU, creates substantial 
heap churn and write amplification, and increases temporary disk pressure.

CEP-66 proposes splitting eligible compressed SSTables by retaining contiguous 
runs of their existing compression chunks. Cassandra would rebuild the child 
SSTables' indexes and other derived components without deserializing, 
serializing, or recompressing their rows.

"Zero-copy" here primarily means reusing the encoded bytes instead of rewriting 
rows. On filesystems that support range reflinks, Cassandra can also share the 
underlying extents meaning no new data written. Other filesystems, including 
ext4, would copy the already-compressed bytes and still avoid the row rewrite.

The proposal is staged. It starts with an opt-in `sstablesplit --zero-copy` 
mode for BIG-format SSTables in Cassandra 7.0/trunk. Later phases add BTI 
support, secondary indexes, anticompaction, and partial-range streaming. 
Existing implementations remain the default and provide the fallback for 
unsupported inputs.

I'd particularly appreciate feedback on:

- The retained-prefix representation and proposed Cassandra 7.0 SSTable format 
change
- Rebuilding or conservatively deriving child metadata without decoding rows
- The integrity and performance tradeoff around `Digest.crc32` generation
- The staged rollout, compatibility rules, and fallback behavior
- Any correctness, operational, or filesystem concerns the proposal has missed

Thanks, and I look forward to the discussion.

Regards,
Chris Lohfink

Reply via email to