Regarding the BTI scenario, if I am not mistaken for narrow partitions we
do not have a full partition key within a primary index, it is only in Data
file, so to rebuild things like bloom filters we will have to read the data
file itself... I suppose it can be only one of the reasons why DSE
implementation does not split  primary indexes.

On Tue, 22 Sept 2026 at 20:03, Patrick McFadin <[email protected]> wrote:

> This is such a clever way to solve a problem I wish would go away. The
> only negative aspect of this proposal is that it isn't available today.
>
> +1 from me.
>
> On Tue, Sep 22, 2026 at 7:18 AM Andy Tolbert <[email protected]> wrote:
>
>> Realized I didn't share a +1 in my previous message, so adding my +1!
>>
>> Thanks,
>> Andy
>>
>> On Tue, Sep 22, 2026, at 2:01 PM, Francisco Guerrero wrote:
>> > +1. Anticompaction is an area that causes a lot of pain and
>> > seeing this proposal makes me hopeful that things will get
>> > better soon.
>> >
>> > Also, I think we should consider bringing this work to 6.0 as
>> > well assuming this lands before we release 6.0
>> >
>> > Best,
>> > - Francisco
>> >
>> > On 2026/09/22 13:43:37 Abe Ratnofsky wrote:
>> >> I’m +1 on the CEP.
>> >>
>> >> Saving bandwidth is meaningful, particularly for those running on
>> networked disks where bandwidth has a lower ceiling and a direct marginal
>> cost. The project currently recommends NVMe but the cost and convenience
>> benefits of newer generations of networked disks are meaningful.
>> >>
>> >> Regarding portability: this feels no different from supporting
>> multiple JDK versions, or recommending ACCP, etc. You’ll get better
>> performance on certain systems where certain features are available. This
>> change also introduces performance improvements for filesystems that do not
>> support FICLONERANGE like ext4.
>> >>
>> >> I do wish it were possible for us to version SSTables in a way that
>> let users experiment with this more easily. Requiring it land in a new
>> major keeps it further away from many users who would benefit.
>> >>
>> >> On Tue, Sep 22, 2026, at 5:38 AM, Benedict Elliott Smith wrote:
>> >> > My biggest concern with this proposal is whether linking inodes is
>> >> > justified. This narrows the applicability (only supported
>> filesystems)
>> >> > and increases the complexity, but in return saves bandwidth and
>> allows
>> >> > it to function when disk space is exhausted.
>> >> >
>> >> > I am of the opinion it is probably not justified, as bandwidth is
>> not I
>> >> > think typically a major constraint - we don't generally saturate the
>> >> > disk because we are CPU-inefficient. Simply copying the files would
>> be
>> >> > a better starting point as much simpler and more portable. We have
>> >> > bigger problems when we are out of space.
>> >> >
>> >> > Another thing to balance is whether this complexity is justified for
>> a
>> >> > stop-gap measure, if we expect this to be made defunct by both
>> >> > Branimir's new file format (which permits cheaper slicing) and
>> mutation
>> >> > tracking (which should eliminate the need for anti-compaction).
>> >> >
>> >> > Separately, I wonder (if we desperately want it in the meantime)
>> >> > whether DataStax are willing to contribute their version of this, if
>> it
>> >> > already exists and is already validated on real workloads.
>> >> >
>> >> > Finally, have we explored simply removing anti-compaction instead?
>> This
>> >> > would require I think a couple of components: 1) per-range repair
>> >> > metadata; 2) either per-range sstable invalidations (so that
>> compaction
>> >> > may proceed on parts of the repaired/unrepaired file independently,
>> >> > permitting it to be replaced in both sets), or rewriting the
>> >> > uncompacted part(s) of the sstable. I may be missing some other
>> >> > complexity, but this seems quite tractable.
>> >> >
>> >> >
>> >> >
>> >> > On 2026/09/16 12:25:28 Chris Lohfink wrote:
>> >> >> Thanks, Branimir. The shared-index approach is an attractive option
>> >> >> especially for transient local operations such as anticompaction.
>> >> >>
>> >> >> The tradeoff appears to be where the complexity lives. Reusing the
>> original
>> >> >> primary index makes slice creation cheaper, but its positions
>> remain in the
>> >> >> parent Data.db coordinate space. Readers and tools must therefore
>> >> >> understand that the Data.db component is a slice and translate those
>> >> >> positions before accessing the local file.
>> >> >>
>> >> >> The current proposal instead pays that cost once during splitting.
>> It
>> >> >> rebases Index.db positions, slices CompressionInfo.db, and rebuilds
>> or
>> >> >> apportions the child’s derived metadata. This preserves the usual
>> idea that
>> >> >> each SSTable is a self-contained artifact whose components share one
>> >> >> coordinate space and lifecycle. Index reconstruction is so cheap its
>> >> >> comparatively free relative to reading or copying Data.db, while
>> also
>> >> >> producing child-specific summaries, Bloom filters, and statistics
>> (100's of
>> >> >> ms range on *huge* sstables)
>> >> >>
>> >> >> For durable outputs, I prefer keeping that complexity at creation
>> time.
>> >> >> Sharing one physical index file among several logical SSTables would
>> >> >> introduce additional lifecycle and bookkeeping cases around
>> deletion,
>> >> >> snapshots, backup and restore, import, and third-party tooling.
>> Independent
>> >> >> files that happen to share filesystem extents are less concerning
>> because
>> >> >> they retain normal component ownership semantics.
>> >> >>
>> >> >> That said, the DataStax approach is worth benchmarking, and it may
>> be a
>> >> >> better fit for BTI or another format where slicing is designed in
>> as a
>> >> >> first-class property.
>> >> >>
>> >> >> > Separately, being able to easily slice and dice files without
>> looking
>> >> >> inside them is a key consideration in the file format we are
>> working on for
>> >> >> CEP-57.
>> >> >>
>> >> >> Agreed,nCEP-57 seems like the right place to make sliceability a
>> native
>> >> >> format property. My goal here is narrower: provide the capability
>> for BIG
>> >> >> SSTables and current formats in the meantime without permanently
>> >> >> introducing slice-coordinate awareness throughout BIG’s read path.
>> BIG will
>> >> >> remain in use for some time, and this mechanism can deliver most of
>> the
>> >> >> benefit with relatively contained changes.
>> >> >>
>> >> >> On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov <[email protected]>
>> wrote:
>> >> >>
>> >> >> > Hello Chris,
>> >> >> >
>> >> >> > A while back we implemented a similar approach for DSE's version
>> of
>> >> >> > zero-copy streaming. The main difference between our approach and
>> yours is
>> >> >> > that we decided not to split the primary index files and instead
>> use them
>> >> >> > as they are, with filtering based on the start and end key of the
>> section.
>> >> >> > In the context of local operations like anticompaction, the
>> latter may be a
>> >> >> > better approach as one can share the index files between all
>> resulting
>> >> >> > sections (and, of course, no index reconstruction is necessary).
>> >> >> >
>> >> >> > This code is not currently part of the Apache codebase, but
>> DataStax's
>> >> >> > open source fork includes support for reading these files, whose
>> >> >> > implementation (commit
>> >> >> >
>> https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422
>> )
>> >> >> > may have some ideas you can use in designing your solution.
>> >> >> >
>> >> >> > Separately, being able to easily slice and dice files without
>> looking
>> >> >> > inside them is a key consideration in the file format we are
>> working on for
>> >> >> > CEP-57.
>> >> >> >
>> >> >> > Regards,
>> >> >> > Branimir
>> >> >> >
>> >> >> > On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella <
>> >> >> > [email protected]> wrote:
>> >> >> >
>> >> >> >> Really nice idea Chris!
>> >> >> >>
>> >> >> >> One thing I think it’s worth flagging for the proposal:
>> >> >> >>
>> >> >> >> this CEP introduces a new SSTable major version. That's a
>> relevant change
>> >> >> >> for anything outside nodetool/the core read path that parses
>> SSTables
>> >> >> >> directly. e.g. analytics library uses the concept of per major
>> version
>> >> >> >> bridge to deserialize.
>> >> >> >>
>> >> >> >> With this approach, current bridges don't have a notion of "data
>> doesn't
>> >> >> >> start at logical offset zero” - they assume a chunk's first
>> partition
>> >> >> >> begins the SSTable's data - which is fine (this is a new
>> version!). It'd
>> >> >> >> help if the CEP explicitly calls out that third-party/off-node
>> SSTable
>> >> >> >> readers are a compatibility surface here, not just in-process
>> Cassandra
>> >> >> >> binaries.
>> >> >> >>
>> >> >> >> Bernardo
>> >> >> >>
>> >> >> >> *From: *Chris Lohfink <[email protected]>
>> >> >> >> *Date: *Monday, 14 September 2026 at 21:23
>> >> >> >> *To: *[email protected] <[email protected]>
>> >> >> >> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting
>> >> >> >>
>> >> >> >> Hi everyone,
>> >> >> >>
>> >> >> >> I'd like to open CEP-66, Zero-copy SSTable splitting, for
>> discussion:
>> >> >> >>
>> >> >> >>
>> >> >> >> *
>> https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting*
>> >> >> >> <
>> https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting
>> >
>> >> >> >>
>> >> >> >>
>> >> >> >> Anticompaction and partial-range streaming currently rewrite
>> rows whose
>> >> >> >> encoded representation already exists on disk. This consumes
>> CPU, creates
>> >> >> >> substantial heap churn and write amplification, and increases
>> temporary
>> >> >> >> disk pressure.
>> >> >> >>
>> >> >> >> CEP-66 proposes splitting eligible compressed SSTables by
>> retaining
>> >> >> >> contiguous runs of their existing compression chunks. Cassandra
>> would
>> >> >> >> rebuild the child SSTables' indexes and other derived components
>> without
>> >> >> >> deserializing, serializing, or recompressing their rows.
>> >> >> >>
>> >> >> >> "Zero-copy" here primarily means reusing the encoded bytes
>> instead of
>> >> >> >> rewriting rows. On filesystems that support range reflinks,
>> Cassandra can
>> >> >> >> also share the underlying extents meaning no new data written.
>> Other
>> >> >> >> filesystems, including ext4, would copy the already-compressed
>> bytes and
>> >> >> >> still avoid the row rewrite.
>> >> >> >>
>> >> >> >> The proposal is staged. It starts with an opt-in `sstablesplit
>> >> >> >> --zero-copy` mode for BIG-format SSTables in Cassandra
>> 7.0/trunk. Later
>> >> >> >> phases add BTI support, secondary indexes, anticompaction, and
>> >> >> >> partial-range streaming. Existing implementations remain the
>> default and
>> >> >> >> provide the fallback for unsupported inputs.
>> >> >> >>
>> >> >> >> I'd particularly appreciate feedback on:
>> >> >> >>
>> >> >> >> - The retained-prefix representation and proposed Cassandra 7.0
>> SSTable
>> >> >> >> format change
>> >> >> >> - Rebuilding or conservatively deriving child metadata without
>> decoding
>> >> >> >> rows
>> >> >> >> - The integrity and performance tradeoff around `Digest.crc32`
>> generation
>> >> >> >> - The staged rollout, compatibility rules, and fallback behavior
>> >> >> >> - Any correctness, operational, or filesystem concerns the
>> proposal has
>> >> >> >> missed
>> >> >> >>
>> >> >> >> Thanks, and I look forward to the discussion.
>> >> >> >>
>> >> >> >> Regards,
>> >> >> >> Chris Lohfink
>> >> >> >>
>> >> >> >
>> >> >>
>> >>
>>
>

-- 
Dmitry Konstantinov

Reply via email to