Incremental repair is critical for database consistency, so resolving the anti-compaction bottleneck proposed in this CEP is a great step forward. Since many Cassandra features aren't fully tested at scale, programmatic guardrails are essential—as long as they are strictly enforced in code, I do not see any issues with the proposal.
Jaydeep On Tue, Sep 15, 2026 at 8:56 AM Jon Haddad <[email protected]> wrote: > This is a great idea. Anti-compaction is a major headache for users > today. > > Jon > > On Mon, Sep 14, 2026 at 12:23 PM Chris Lohfink <[email protected]> > wrote: > >> Hi everyone, >> >> I'd like to open CEP-66, Zero-copy SSTable splitting, for discussion: >> >> *https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting >> <https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting> >> * >> >> Anticompaction and partial-range streaming currently rewrite rows whose >> encoded representation already exists on disk. This consumes CPU, creates >> substantial heap churn and write amplification, and increases temporary >> disk pressure. >> >> CEP-66 proposes splitting eligible compressed SSTables by retaining >> contiguous runs of their existing compression chunks. Cassandra would >> rebuild the child SSTables' indexes and other derived components without >> deserializing, serializing, or recompressing their rows. >> >> "Zero-copy" here primarily means reusing the encoded bytes instead of >> rewriting rows. On filesystems that support range reflinks, Cassandra can >> also share the underlying extents meaning no new data written. Other >> filesystems, including ext4, would copy the already-compressed bytes and >> still avoid the row rewrite. >> >> The proposal is staged. It starts with an opt-in `sstablesplit >> --zero-copy` mode for BIG-format SSTables in Cassandra 7.0/trunk. Later >> phases add BTI support, secondary indexes, anticompaction, and >> partial-range streaming. Existing implementations remain the default and >> provide the fallback for unsupported inputs. >> >> I'd particularly appreciate feedback on: >> >> - The retained-prefix representation and proposed Cassandra 7.0 SSTable >> format change >> - Rebuilding or conservatively deriving child metadata without decoding >> rows >> - The integrity and performance tradeoff around `Digest.crc32` generation >> - The staged rollout, compatibility rules, and fallback behavior >> - Any correctness, operational, or filesystem concerns the proposal has >> missed >> >> Thanks, and I look forward to the discussion. >> >> Regards, >> Chris Lohfink >> >
