Yes, you don't need to deserialize everything, only the partition key, but you still have to decompress the entire chunk, so the overhead could be quite significant.
I don't think this is a blocker for starting with an implementation, but it’s probably worth keeping in mind when designing the overall logic. On Wed, 30 Sept 2026 at 23:34, Chris Lohfink <[email protected]> wrote: > I actually believe you can still do it in BTI; it just requires reading > some of the Data component. We wont need to deserialize everything like > current but will still have IO costs. I think when we get to the BTI part > of implementation we can experiment with a few different approaches more > thoroughly including possibly the DSE implementation but I'm hesitant to > have sstables be dependent on each other's components for operational > simplicity. There's also the option of reusing the parents bloom filter at > the cost of increased false positives. > > Chris > > On Wed, Sep 30, 2026 at 5:11 PM Dmitry Konstantinov <[email protected]> > wrote: > >> Regarding the BTI scenario, if I am not mistaken for narrow partitions we >> do not have a full partition key within a primary index, it is only in Data >> file, so to rebuild things like bloom filters we will have to read the data >> file itself... I suppose it can be only one of the reasons why DSE >> implementation does not split primary indexes. >> >> On Tue, 22 Sept 2026 at 20:03, Patrick McFadin <[email protected]> >> wrote: >> >>> This is such a clever way to solve a problem I wish would go away. The >>> only negative aspect of this proposal is that it isn't available today. >>> >>> +1 from me. >>> >>> On Tue, Sep 22, 2026 at 7:18 AM Andy Tolbert <[email protected]> >>> wrote: >>> >>>> Realized I didn't share a +1 in my previous message, so adding my +1! >>>> >>>> Thanks, >>>> Andy >>>> >>>> On Tue, Sep 22, 2026, at 2:01 PM, Francisco Guerrero wrote: >>>> > +1. Anticompaction is an area that causes a lot of pain and >>>> > seeing this proposal makes me hopeful that things will get >>>> > better soon. >>>> > >>>> > Also, I think we should consider bringing this work to 6.0 as >>>> > well assuming this lands before we release 6.0 >>>> > >>>> > Best, >>>> > - Francisco >>>> > >>>> > On 2026/09/22 13:43:37 Abe Ratnofsky wrote: >>>> >> I’m +1 on the CEP. >>>> >> >>>> >> Saving bandwidth is meaningful, particularly for those running on >>>> networked disks where bandwidth has a lower ceiling and a direct marginal >>>> cost. The project currently recommends NVMe but the cost and convenience >>>> benefits of newer generations of networked disks are meaningful. >>>> >> >>>> >> Regarding portability: this feels no different from supporting >>>> multiple JDK versions, or recommending ACCP, etc. You’ll get better >>>> performance on certain systems where certain features are available. This >>>> change also introduces performance improvements for filesystems that do not >>>> support FICLONERANGE like ext4. >>>> >> >>>> >> I do wish it were possible for us to version SSTables in a way that >>>> let users experiment with this more easily. Requiring it land in a new >>>> major keeps it further away from many users who would benefit. >>>> >> >>>> >> On Tue, Sep 22, 2026, at 5:38 AM, Benedict Elliott Smith wrote: >>>> >> > My biggest concern with this proposal is whether linking inodes is >>>> >> > justified. This narrows the applicability (only supported >>>> filesystems) >>>> >> > and increases the complexity, but in return saves bandwidth and >>>> allows >>>> >> > it to function when disk space is exhausted. >>>> >> > >>>> >> > I am of the opinion it is probably not justified, as bandwidth is >>>> not I >>>> >> > think typically a major constraint - we don't generally saturate >>>> the >>>> >> > disk because we are CPU-inefficient. Simply copying the files >>>> would be >>>> >> > a better starting point as much simpler and more portable. We have >>>> >> > bigger problems when we are out of space. >>>> >> > >>>> >> > Another thing to balance is whether this complexity is justified >>>> for a >>>> >> > stop-gap measure, if we expect this to be made defunct by both >>>> >> > Branimir's new file format (which permits cheaper slicing) and >>>> mutation >>>> >> > tracking (which should eliminate the need for anti-compaction). >>>> >> > >>>> >> > Separately, I wonder (if we desperately want it in the meantime) >>>> >> > whether DataStax are willing to contribute their version of this, >>>> if it >>>> >> > already exists and is already validated on real workloads. >>>> >> > >>>> >> > Finally, have we explored simply removing anti-compaction instead? >>>> This >>>> >> > would require I think a couple of components: 1) per-range repair >>>> >> > metadata; 2) either per-range sstable invalidations (so that >>>> compaction >>>> >> > may proceed on parts of the repaired/unrepaired file >>>> independently, >>>> >> > permitting it to be replaced in both sets), or rewriting the >>>> >> > uncompacted part(s) of the sstable. I may be missing some other >>>> >> > complexity, but this seems quite tractable. >>>> >> > >>>> >> > >>>> >> > >>>> >> > On 2026/09/16 12:25:28 Chris Lohfink wrote: >>>> >> >> Thanks, Branimir. The shared-index approach is an attractive >>>> option >>>> >> >> especially for transient local operations such as anticompaction. >>>> >> >> >>>> >> >> The tradeoff appears to be where the complexity lives. Reusing >>>> the original >>>> >> >> primary index makes slice creation cheaper, but its positions >>>> remain in the >>>> >> >> parent Data.db coordinate space. Readers and tools must therefore >>>> >> >> understand that the Data.db component is a slice and translate >>>> those >>>> >> >> positions before accessing the local file. >>>> >> >> >>>> >> >> The current proposal instead pays that cost once during >>>> splitting. It >>>> >> >> rebases Index.db positions, slices CompressionInfo.db, and >>>> rebuilds or >>>> >> >> apportions the child’s derived metadata. This preserves the usual >>>> idea that >>>> >> >> each SSTable is a self-contained artifact whose components share >>>> one >>>> >> >> coordinate space and lifecycle. Index reconstruction is so cheap >>>> its >>>> >> >> comparatively free relative to reading or copying Data.db, while >>>> also >>>> >> >> producing child-specific summaries, Bloom filters, and statistics >>>> (100's of >>>> >> >> ms range on *huge* sstables) >>>> >> >> >>>> >> >> For durable outputs, I prefer keeping that complexity at creation >>>> time. >>>> >> >> Sharing one physical index file among several logical SSTables >>>> would >>>> >> >> introduce additional lifecycle and bookkeeping cases around >>>> deletion, >>>> >> >> snapshots, backup and restore, import, and third-party tooling. >>>> Independent >>>> >> >> files that happen to share filesystem extents are less concerning >>>> because >>>> >> >> they retain normal component ownership semantics. >>>> >> >> >>>> >> >> That said, the DataStax approach is worth benchmarking, and it >>>> may be a >>>> >> >> better fit for BTI or another format where slicing is designed in >>>> as a >>>> >> >> first-class property. >>>> >> >> >>>> >> >> > Separately, being able to easily slice and dice files without >>>> looking >>>> >> >> inside them is a key consideration in the file format we are >>>> working on for >>>> >> >> CEP-57. >>>> >> >> >>>> >> >> Agreed,nCEP-57 seems like the right place to make sliceability a >>>> native >>>> >> >> format property. My goal here is narrower: provide the capability >>>> for BIG >>>> >> >> SSTables and current formats in the meantime without permanently >>>> >> >> introducing slice-coordinate awareness throughout BIG’s read >>>> path. BIG will >>>> >> >> remain in use for some time, and this mechanism can deliver most >>>> of the >>>> >> >> benefit with relatively contained changes. >>>> >> >> >>>> >> >> On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov < >>>> [email protected]> wrote: >>>> >> >> >>>> >> >> > Hello Chris, >>>> >> >> > >>>> >> >> > A while back we implemented a similar approach for DSE's >>>> version of >>>> >> >> > zero-copy streaming. The main difference between our approach >>>> and yours is >>>> >> >> > that we decided not to split the primary index files and >>>> instead use them >>>> >> >> > as they are, with filtering based on the start and end key of >>>> the section. >>>> >> >> > In the context of local operations like anticompaction, the >>>> latter may be a >>>> >> >> > better approach as one can share the index files between all >>>> resulting >>>> >> >> > sections (and, of course, no index reconstruction is necessary). >>>> >> >> > >>>> >> >> > This code is not currently part of the Apache codebase, but >>>> DataStax's >>>> >> >> > open source fork includes support for reading these files, whose >>>> >> >> > implementation (commit >>>> >> >> > >>>> https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422 >>>> ) >>>> >> >> > may have some ideas you can use in designing your solution. >>>> >> >> > >>>> >> >> > Separately, being able to easily slice and dice files without >>>> looking >>>> >> >> > inside them is a key consideration in the file format we are >>>> working on for >>>> >> >> > CEP-57. >>>> >> >> > >>>> >> >> > Regards, >>>> >> >> > Branimir >>>> >> >> > >>>> >> >> > On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella < >>>> >> >> > [email protected]> wrote: >>>> >> >> > >>>> >> >> >> Really nice idea Chris! >>>> >> >> >> >>>> >> >> >> One thing I think it’s worth flagging for the proposal: >>>> >> >> >> >>>> >> >> >> this CEP introduces a new SSTable major version. That's a >>>> relevant change >>>> >> >> >> for anything outside nodetool/the core read path that parses >>>> SSTables >>>> >> >> >> directly. e.g. analytics library uses the concept of per major >>>> version >>>> >> >> >> bridge to deserialize. >>>> >> >> >> >>>> >> >> >> With this approach, current bridges don't have a notion of >>>> "data doesn't >>>> >> >> >> start at logical offset zero” - they assume a chunk's first >>>> partition >>>> >> >> >> begins the SSTable's data - which is fine (this is a new >>>> version!). It'd >>>> >> >> >> help if the CEP explicitly calls out that third-party/off-node >>>> SSTable >>>> >> >> >> readers are a compatibility surface here, not just in-process >>>> Cassandra >>>> >> >> >> binaries. >>>> >> >> >> >>>> >> >> >> Bernardo >>>> >> >> >> >>>> >> >> >> *From: *Chris Lohfink <[email protected]> >>>> >> >> >> *Date: *Monday, 14 September 2026 at 21:23 >>>> >> >> >> *To: *[email protected] <[email protected]> >>>> >> >> >> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting >>>> >> >> >> >>>> >> >> >> Hi everyone, >>>> >> >> >> >>>> >> >> >> I'd like to open CEP-66, Zero-copy SSTable splitting, for >>>> discussion: >>>> >> >> >> >>>> >> >> >> >>>> >> >> >> * >>>> https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting* >>>> >> >> >> < >>>> https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting >>>> > >>>> >> >> >> >>>> >> >> >> >>>> >> >> >> Anticompaction and partial-range streaming currently rewrite >>>> rows whose >>>> >> >> >> encoded representation already exists on disk. This consumes >>>> CPU, creates >>>> >> >> >> substantial heap churn and write amplification, and increases >>>> temporary >>>> >> >> >> disk pressure. >>>> >> >> >> >>>> >> >> >> CEP-66 proposes splitting eligible compressed SSTables by >>>> retaining >>>> >> >> >> contiguous runs of their existing compression chunks. >>>> Cassandra would >>>> >> >> >> rebuild the child SSTables' indexes and other derived >>>> components without >>>> >> >> >> deserializing, serializing, or recompressing their rows. >>>> >> >> >> >>>> >> >> >> "Zero-copy" here primarily means reusing the encoded bytes >>>> instead of >>>> >> >> >> rewriting rows. On filesystems that support range reflinks, >>>> Cassandra can >>>> >> >> >> also share the underlying extents meaning no new data written. >>>> Other >>>> >> >> >> filesystems, including ext4, would copy the already-compressed >>>> bytes and >>>> >> >> >> still avoid the row rewrite. >>>> >> >> >> >>>> >> >> >> The proposal is staged. It starts with an opt-in `sstablesplit >>>> >> >> >> --zero-copy` mode for BIG-format SSTables in Cassandra >>>> 7.0/trunk. Later >>>> >> >> >> phases add BTI support, secondary indexes, anticompaction, and >>>> >> >> >> partial-range streaming. Existing implementations remain the >>>> default and >>>> >> >> >> provide the fallback for unsupported inputs. >>>> >> >> >> >>>> >> >> >> I'd particularly appreciate feedback on: >>>> >> >> >> >>>> >> >> >> - The retained-prefix representation and proposed Cassandra >>>> 7.0 SSTable >>>> >> >> >> format change >>>> >> >> >> - Rebuilding or conservatively deriving child metadata without >>>> decoding >>>> >> >> >> rows >>>> >> >> >> - The integrity and performance tradeoff around `Digest.crc32` >>>> generation >>>> >> >> >> - The staged rollout, compatibility rules, and fallback >>>> behavior >>>> >> >> >> - Any correctness, operational, or filesystem concerns the >>>> proposal has >>>> >> >> >> missed >>>> >> >> >> >>>> >> >> >> Thanks, and I look forward to the discussion. >>>> >> >> >> >>>> >> >> >> Regards, >>>> >> >> >> Chris Lohfink >>>> >> >> >> >>>> >> >> > >>>> >> >> >>>> >> >>>> >>> >> >> -- >> Dmitry Konstantinov >> > -- Dmitry Konstantinov
