Regarding the BTI scenario, if I am not mistaken for narrow partitions we do not have a full partition key within a primary index, it is only in Data file, so to rebuild things like bloom filters we will have to read the data file itself... I suppose it can be only one of the reasons why DSE implementation does not split primary indexes.
On Tue, 22 Sept 2026 at 20:03, Patrick McFadin <[email protected]> wrote: > This is such a clever way to solve a problem I wish would go away. The > only negative aspect of this proposal is that it isn't available today. > > +1 from me. > > On Tue, Sep 22, 2026 at 7:18 AM Andy Tolbert <[email protected]> wrote: > >> Realized I didn't share a +1 in my previous message, so adding my +1! >> >> Thanks, >> Andy >> >> On Tue, Sep 22, 2026, at 2:01 PM, Francisco Guerrero wrote: >> > +1. Anticompaction is an area that causes a lot of pain and >> > seeing this proposal makes me hopeful that things will get >> > better soon. >> > >> > Also, I think we should consider bringing this work to 6.0 as >> > well assuming this lands before we release 6.0 >> > >> > Best, >> > - Francisco >> > >> > On 2026/09/22 13:43:37 Abe Ratnofsky wrote: >> >> I’m +1 on the CEP. >> >> >> >> Saving bandwidth is meaningful, particularly for those running on >> networked disks where bandwidth has a lower ceiling and a direct marginal >> cost. The project currently recommends NVMe but the cost and convenience >> benefits of newer generations of networked disks are meaningful. >> >> >> >> Regarding portability: this feels no different from supporting >> multiple JDK versions, or recommending ACCP, etc. You’ll get better >> performance on certain systems where certain features are available. This >> change also introduces performance improvements for filesystems that do not >> support FICLONERANGE like ext4. >> >> >> >> I do wish it were possible for us to version SSTables in a way that >> let users experiment with this more easily. Requiring it land in a new >> major keeps it further away from many users who would benefit. >> >> >> >> On Tue, Sep 22, 2026, at 5:38 AM, Benedict Elliott Smith wrote: >> >> > My biggest concern with this proposal is whether linking inodes is >> >> > justified. This narrows the applicability (only supported >> filesystems) >> >> > and increases the complexity, but in return saves bandwidth and >> allows >> >> > it to function when disk space is exhausted. >> >> > >> >> > I am of the opinion it is probably not justified, as bandwidth is >> not I >> >> > think typically a major constraint - we don't generally saturate the >> >> > disk because we are CPU-inefficient. Simply copying the files would >> be >> >> > a better starting point as much simpler and more portable. We have >> >> > bigger problems when we are out of space. >> >> > >> >> > Another thing to balance is whether this complexity is justified for >> a >> >> > stop-gap measure, if we expect this to be made defunct by both >> >> > Branimir's new file format (which permits cheaper slicing) and >> mutation >> >> > tracking (which should eliminate the need for anti-compaction). >> >> > >> >> > Separately, I wonder (if we desperately want it in the meantime) >> >> > whether DataStax are willing to contribute their version of this, if >> it >> >> > already exists and is already validated on real workloads. >> >> > >> >> > Finally, have we explored simply removing anti-compaction instead? >> This >> >> > would require I think a couple of components: 1) per-range repair >> >> > metadata; 2) either per-range sstable invalidations (so that >> compaction >> >> > may proceed on parts of the repaired/unrepaired file independently, >> >> > permitting it to be replaced in both sets), or rewriting the >> >> > uncompacted part(s) of the sstable. I may be missing some other >> >> > complexity, but this seems quite tractable. >> >> > >> >> > >> >> > >> >> > On 2026/09/16 12:25:28 Chris Lohfink wrote: >> >> >> Thanks, Branimir. The shared-index approach is an attractive option >> >> >> especially for transient local operations such as anticompaction. >> >> >> >> >> >> The tradeoff appears to be where the complexity lives. Reusing the >> original >> >> >> primary index makes slice creation cheaper, but its positions >> remain in the >> >> >> parent Data.db coordinate space. Readers and tools must therefore >> >> >> understand that the Data.db component is a slice and translate those >> >> >> positions before accessing the local file. >> >> >> >> >> >> The current proposal instead pays that cost once during splitting. >> It >> >> >> rebases Index.db positions, slices CompressionInfo.db, and rebuilds >> or >> >> >> apportions the child’s derived metadata. This preserves the usual >> idea that >> >> >> each SSTable is a self-contained artifact whose components share one >> >> >> coordinate space and lifecycle. Index reconstruction is so cheap its >> >> >> comparatively free relative to reading or copying Data.db, while >> also >> >> >> producing child-specific summaries, Bloom filters, and statistics >> (100's of >> >> >> ms range on *huge* sstables) >> >> >> >> >> >> For durable outputs, I prefer keeping that complexity at creation >> time. >> >> >> Sharing one physical index file among several logical SSTables would >> >> >> introduce additional lifecycle and bookkeeping cases around >> deletion, >> >> >> snapshots, backup and restore, import, and third-party tooling. >> Independent >> >> >> files that happen to share filesystem extents are less concerning >> because >> >> >> they retain normal component ownership semantics. >> >> >> >> >> >> That said, the DataStax approach is worth benchmarking, and it may >> be a >> >> >> better fit for BTI or another format where slicing is designed in >> as a >> >> >> first-class property. >> >> >> >> >> >> > Separately, being able to easily slice and dice files without >> looking >> >> >> inside them is a key consideration in the file format we are >> working on for >> >> >> CEP-57. >> >> >> >> >> >> Agreed,nCEP-57 seems like the right place to make sliceability a >> native >> >> >> format property. My goal here is narrower: provide the capability >> for BIG >> >> >> SSTables and current formats in the meantime without permanently >> >> >> introducing slice-coordinate awareness throughout BIG’s read path. >> BIG will >> >> >> remain in use for some time, and this mechanism can deliver most of >> the >> >> >> benefit with relatively contained changes. >> >> >> >> >> >> On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov <[email protected]> >> wrote: >> >> >> >> >> >> > Hello Chris, >> >> >> > >> >> >> > A while back we implemented a similar approach for DSE's version >> of >> >> >> > zero-copy streaming. The main difference between our approach and >> yours is >> >> >> > that we decided not to split the primary index files and instead >> use them >> >> >> > as they are, with filtering based on the start and end key of the >> section. >> >> >> > In the context of local operations like anticompaction, the >> latter may be a >> >> >> > better approach as one can share the index files between all >> resulting >> >> >> > sections (and, of course, no index reconstruction is necessary). >> >> >> > >> >> >> > This code is not currently part of the Apache codebase, but >> DataStax's >> >> >> > open source fork includes support for reading these files, whose >> >> >> > implementation (commit >> >> >> > >> https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422 >> ) >> >> >> > may have some ideas you can use in designing your solution. >> >> >> > >> >> >> > Separately, being able to easily slice and dice files without >> looking >> >> >> > inside them is a key consideration in the file format we are >> working on for >> >> >> > CEP-57. >> >> >> > >> >> >> > Regards, >> >> >> > Branimir >> >> >> > >> >> >> > On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella < >> >> >> > [email protected]> wrote: >> >> >> > >> >> >> >> Really nice idea Chris! >> >> >> >> >> >> >> >> One thing I think it’s worth flagging for the proposal: >> >> >> >> >> >> >> >> this CEP introduces a new SSTable major version. That's a >> relevant change >> >> >> >> for anything outside nodetool/the core read path that parses >> SSTables >> >> >> >> directly. e.g. analytics library uses the concept of per major >> version >> >> >> >> bridge to deserialize. >> >> >> >> >> >> >> >> With this approach, current bridges don't have a notion of "data >> doesn't >> >> >> >> start at logical offset zero” - they assume a chunk's first >> partition >> >> >> >> begins the SSTable's data - which is fine (this is a new >> version!). It'd >> >> >> >> help if the CEP explicitly calls out that third-party/off-node >> SSTable >> >> >> >> readers are a compatibility surface here, not just in-process >> Cassandra >> >> >> >> binaries. >> >> >> >> >> >> >> >> Bernardo >> >> >> >> >> >> >> >> *From: *Chris Lohfink <[email protected]> >> >> >> >> *Date: *Monday, 14 September 2026 at 21:23 >> >> >> >> *To: *[email protected] <[email protected]> >> >> >> >> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting >> >> >> >> >> >> >> >> Hi everyone, >> >> >> >> >> >> >> >> I'd like to open CEP-66, Zero-copy SSTable splitting, for >> discussion: >> >> >> >> >> >> >> >> >> >> >> >> * >> https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting* >> >> >> >> < >> https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting >> > >> >> >> >> >> >> >> >> >> >> >> >> Anticompaction and partial-range streaming currently rewrite >> rows whose >> >> >> >> encoded representation already exists on disk. This consumes >> CPU, creates >> >> >> >> substantial heap churn and write amplification, and increases >> temporary >> >> >> >> disk pressure. >> >> >> >> >> >> >> >> CEP-66 proposes splitting eligible compressed SSTables by >> retaining >> >> >> >> contiguous runs of their existing compression chunks. Cassandra >> would >> >> >> >> rebuild the child SSTables' indexes and other derived components >> without >> >> >> >> deserializing, serializing, or recompressing their rows. >> >> >> >> >> >> >> >> "Zero-copy" here primarily means reusing the encoded bytes >> instead of >> >> >> >> rewriting rows. On filesystems that support range reflinks, >> Cassandra can >> >> >> >> also share the underlying extents meaning no new data written. >> Other >> >> >> >> filesystems, including ext4, would copy the already-compressed >> bytes and >> >> >> >> still avoid the row rewrite. >> >> >> >> >> >> >> >> The proposal is staged. It starts with an opt-in `sstablesplit >> >> >> >> --zero-copy` mode for BIG-format SSTables in Cassandra >> 7.0/trunk. Later >> >> >> >> phases add BTI support, secondary indexes, anticompaction, and >> >> >> >> partial-range streaming. Existing implementations remain the >> default and >> >> >> >> provide the fallback for unsupported inputs. >> >> >> >> >> >> >> >> I'd particularly appreciate feedback on: >> >> >> >> >> >> >> >> - The retained-prefix representation and proposed Cassandra 7.0 >> SSTable >> >> >> >> format change >> >> >> >> - Rebuilding or conservatively deriving child metadata without >> decoding >> >> >> >> rows >> >> >> >> - The integrity and performance tradeoff around `Digest.crc32` >> generation >> >> >> >> - The staged rollout, compatibility rules, and fallback behavior >> >> >> >> - Any correctness, operational, or filesystem concerns the >> proposal has >> >> >> >> missed >> >> >> >> >> >> >> >> Thanks, and I look forward to the discussion. >> >> >> >> >> >> >> >> Regards, >> >> >> >> Chris Lohfink >> >> >> >> >> >> >> > >> >> >> >> >> >> > -- Dmitry Konstantinov
