Yes, you don't need to deserialize everything, only the partition key, but
you still have to decompress the entire chunk, so the overhead could be
quite significant.

I don't think this is a blocker for starting with an implementation, but
it’s probably worth keeping in mind when designing the overall logic.


On Wed, 30 Sept 2026 at 23:34, Chris Lohfink <[email protected]> wrote:

> I actually believe you can still do it in BTI; it just requires reading
> some of the Data component. We wont need to deserialize everything like
> current but will still have IO costs. I think when we get to the BTI part
> of implementation we can experiment with a few different approaches more
> thoroughly including possibly the DSE implementation but I'm hesitant to
> have sstables be dependent on each other's components for operational
> simplicity. There's also the option of reusing the parents bloom filter at
> the cost of increased false positives.
>
> Chris
>
> On Wed, Sep 30, 2026 at 5:11 PM Dmitry Konstantinov <[email protected]>
> wrote:
>
>> Regarding the BTI scenario, if I am not mistaken for narrow partitions we
>> do not have a full partition key within a primary index, it is only in Data
>> file, so to rebuild things like bloom filters we will have to read the data
>> file itself... I suppose it can be only one of the reasons why DSE
>> implementation does not split  primary indexes.
>>
>> On Tue, 22 Sept 2026 at 20:03, Patrick McFadin <[email protected]>
>> wrote:
>>
>>> This is such a clever way to solve a problem I wish would go away. The
>>> only negative aspect of this proposal is that it isn't available today.
>>>
>>> +1 from me.
>>>
>>> On Tue, Sep 22, 2026 at 7:18 AM Andy Tolbert <[email protected]>
>>> wrote:
>>>
>>>> Realized I didn't share a +1 in my previous message, so adding my +1!
>>>>
>>>> Thanks,
>>>> Andy
>>>>
>>>> On Tue, Sep 22, 2026, at 2:01 PM, Francisco Guerrero wrote:
>>>> > +1. Anticompaction is an area that causes a lot of pain and
>>>> > seeing this proposal makes me hopeful that things will get
>>>> > better soon.
>>>> >
>>>> > Also, I think we should consider bringing this work to 6.0 as
>>>> > well assuming this lands before we release 6.0
>>>> >
>>>> > Best,
>>>> > - Francisco
>>>> >
>>>> > On 2026/09/22 13:43:37 Abe Ratnofsky wrote:
>>>> >> I’m +1 on the CEP.
>>>> >>
>>>> >> Saving bandwidth is meaningful, particularly for those running on
>>>> networked disks where bandwidth has a lower ceiling and a direct marginal
>>>> cost. The project currently recommends NVMe but the cost and convenience
>>>> benefits of newer generations of networked disks are meaningful.
>>>> >>
>>>> >> Regarding portability: this feels no different from supporting
>>>> multiple JDK versions, or recommending ACCP, etc. You’ll get better
>>>> performance on certain systems where certain features are available. This
>>>> change also introduces performance improvements for filesystems that do not
>>>> support FICLONERANGE like ext4.
>>>> >>
>>>> >> I do wish it were possible for us to version SSTables in a way that
>>>> let users experiment with this more easily. Requiring it land in a new
>>>> major keeps it further away from many users who would benefit.
>>>> >>
>>>> >> On Tue, Sep 22, 2026, at 5:38 AM, Benedict Elliott Smith wrote:
>>>> >> > My biggest concern with this proposal is whether linking inodes is
>>>> >> > justified. This narrows the applicability (only supported
>>>> filesystems)
>>>> >> > and increases the complexity, but in return saves bandwidth and
>>>> allows
>>>> >> > it to function when disk space is exhausted.
>>>> >> >
>>>> >> > I am of the opinion it is probably not justified, as bandwidth is
>>>> not I
>>>> >> > think typically a major constraint - we don't generally saturate
>>>> the
>>>> >> > disk because we are CPU-inefficient. Simply copying the files
>>>> would be
>>>> >> > a better starting point as much simpler and more portable. We have
>>>> >> > bigger problems when we are out of space.
>>>> >> >
>>>> >> > Another thing to balance is whether this complexity is justified
>>>> for a
>>>> >> > stop-gap measure, if we expect this to be made defunct by both
>>>> >> > Branimir's new file format (which permits cheaper slicing) and
>>>> mutation
>>>> >> > tracking (which should eliminate the need for anti-compaction).
>>>> >> >
>>>> >> > Separately, I wonder (if we desperately want it in the meantime)
>>>> >> > whether DataStax are willing to contribute their version of this,
>>>> if it
>>>> >> > already exists and is already validated on real workloads.
>>>> >> >
>>>> >> > Finally, have we explored simply removing anti-compaction instead?
>>>> This
>>>> >> > would require I think a couple of components: 1) per-range repair
>>>> >> > metadata; 2) either per-range sstable invalidations (so that
>>>> compaction
>>>> >> > may proceed on parts of the repaired/unrepaired file
>>>> independently,
>>>> >> > permitting it to be replaced in both sets), or rewriting the
>>>> >> > uncompacted part(s) of the sstable. I may be missing some other
>>>> >> > complexity, but this seems quite tractable.
>>>> >> >
>>>> >> >
>>>> >> >
>>>> >> > On 2026/09/16 12:25:28 Chris Lohfink wrote:
>>>> >> >> Thanks, Branimir. The shared-index approach is an attractive
>>>> option
>>>> >> >> especially for transient local operations such as anticompaction.
>>>> >> >>
>>>> >> >> The tradeoff appears to be where the complexity lives. Reusing
>>>> the original
>>>> >> >> primary index makes slice creation cheaper, but its positions
>>>> remain in the
>>>> >> >> parent Data.db coordinate space. Readers and tools must therefore
>>>> >> >> understand that the Data.db component is a slice and translate
>>>> those
>>>> >> >> positions before accessing the local file.
>>>> >> >>
>>>> >> >> The current proposal instead pays that cost once during
>>>> splitting. It
>>>> >> >> rebases Index.db positions, slices CompressionInfo.db, and
>>>> rebuilds or
>>>> >> >> apportions the child’s derived metadata. This preserves the usual
>>>> idea that
>>>> >> >> each SSTable is a self-contained artifact whose components share
>>>> one
>>>> >> >> coordinate space and lifecycle. Index reconstruction is so cheap
>>>> its
>>>> >> >> comparatively free relative to reading or copying Data.db, while
>>>> also
>>>> >> >> producing child-specific summaries, Bloom filters, and statistics
>>>> (100's of
>>>> >> >> ms range on *huge* sstables)
>>>> >> >>
>>>> >> >> For durable outputs, I prefer keeping that complexity at creation
>>>> time.
>>>> >> >> Sharing one physical index file among several logical SSTables
>>>> would
>>>> >> >> introduce additional lifecycle and bookkeeping cases around
>>>> deletion,
>>>> >> >> snapshots, backup and restore, import, and third-party tooling.
>>>> Independent
>>>> >> >> files that happen to share filesystem extents are less concerning
>>>> because
>>>> >> >> they retain normal component ownership semantics.
>>>> >> >>
>>>> >> >> That said, the DataStax approach is worth benchmarking, and it
>>>> may be a
>>>> >> >> better fit for BTI or another format where slicing is designed in
>>>> as a
>>>> >> >> first-class property.
>>>> >> >>
>>>> >> >> > Separately, being able to easily slice and dice files without
>>>> looking
>>>> >> >> inside them is a key consideration in the file format we are
>>>> working on for
>>>> >> >> CEP-57.
>>>> >> >>
>>>> >> >> Agreed,nCEP-57 seems like the right place to make sliceability a
>>>> native
>>>> >> >> format property. My goal here is narrower: provide the capability
>>>> for BIG
>>>> >> >> SSTables and current formats in the meantime without permanently
>>>> >> >> introducing slice-coordinate awareness throughout BIG’s read
>>>> path. BIG will
>>>> >> >> remain in use for some time, and this mechanism can deliver most
>>>> of the
>>>> >> >> benefit with relatively contained changes.
>>>> >> >>
>>>> >> >> On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov <
>>>> [email protected]> wrote:
>>>> >> >>
>>>> >> >> > Hello Chris,
>>>> >> >> >
>>>> >> >> > A while back we implemented a similar approach for DSE's
>>>> version of
>>>> >> >> > zero-copy streaming. The main difference between our approach
>>>> and yours is
>>>> >> >> > that we decided not to split the primary index files and
>>>> instead use them
>>>> >> >> > as they are, with filtering based on the start and end key of
>>>> the section.
>>>> >> >> > In the context of local operations like anticompaction, the
>>>> latter may be a
>>>> >> >> > better approach as one can share the index files between all
>>>> resulting
>>>> >> >> > sections (and, of course, no index reconstruction is necessary).
>>>> >> >> >
>>>> >> >> > This code is not currently part of the Apache codebase, but
>>>> DataStax's
>>>> >> >> > open source fork includes support for reading these files, whose
>>>> >> >> > implementation (commit
>>>> >> >> >
>>>> https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422
>>>> )
>>>> >> >> > may have some ideas you can use in designing your solution.
>>>> >> >> >
>>>> >> >> > Separately, being able to easily slice and dice files without
>>>> looking
>>>> >> >> > inside them is a key consideration in the file format we are
>>>> working on for
>>>> >> >> > CEP-57.
>>>> >> >> >
>>>> >> >> > Regards,
>>>> >> >> > Branimir
>>>> >> >> >
>>>> >> >> > On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella <
>>>> >> >> > [email protected]> wrote:
>>>> >> >> >
>>>> >> >> >> Really nice idea Chris!
>>>> >> >> >>
>>>> >> >> >> One thing I think it’s worth flagging for the proposal:
>>>> >> >> >>
>>>> >> >> >> this CEP introduces a new SSTable major version. That's a
>>>> relevant change
>>>> >> >> >> for anything outside nodetool/the core read path that parses
>>>> SSTables
>>>> >> >> >> directly. e.g. analytics library uses the concept of per major
>>>> version
>>>> >> >> >> bridge to deserialize.
>>>> >> >> >>
>>>> >> >> >> With this approach, current bridges don't have a notion of
>>>> "data doesn't
>>>> >> >> >> start at logical offset zero” - they assume a chunk's first
>>>> partition
>>>> >> >> >> begins the SSTable's data - which is fine (this is a new
>>>> version!). It'd
>>>> >> >> >> help if the CEP explicitly calls out that third-party/off-node
>>>> SSTable
>>>> >> >> >> readers are a compatibility surface here, not just in-process
>>>> Cassandra
>>>> >> >> >> binaries.
>>>> >> >> >>
>>>> >> >> >> Bernardo
>>>> >> >> >>
>>>> >> >> >> *From: *Chris Lohfink <[email protected]>
>>>> >> >> >> *Date: *Monday, 14 September 2026 at 21:23
>>>> >> >> >> *To: *[email protected] <[email protected]>
>>>> >> >> >> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting
>>>> >> >> >>
>>>> >> >> >> Hi everyone,
>>>> >> >> >>
>>>> >> >> >> I'd like to open CEP-66, Zero-copy SSTable splitting, for
>>>> discussion:
>>>> >> >> >>
>>>> >> >> >>
>>>> >> >> >> *
>>>> https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting*
>>>> >> >> >> <
>>>> https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting
>>>> >
>>>> >> >> >>
>>>> >> >> >>
>>>> >> >> >> Anticompaction and partial-range streaming currently rewrite
>>>> rows whose
>>>> >> >> >> encoded representation already exists on disk. This consumes
>>>> CPU, creates
>>>> >> >> >> substantial heap churn and write amplification, and increases
>>>> temporary
>>>> >> >> >> disk pressure.
>>>> >> >> >>
>>>> >> >> >> CEP-66 proposes splitting eligible compressed SSTables by
>>>> retaining
>>>> >> >> >> contiguous runs of their existing compression chunks.
>>>> Cassandra would
>>>> >> >> >> rebuild the child SSTables' indexes and other derived
>>>> components without
>>>> >> >> >> deserializing, serializing, or recompressing their rows.
>>>> >> >> >>
>>>> >> >> >> "Zero-copy" here primarily means reusing the encoded bytes
>>>> instead of
>>>> >> >> >> rewriting rows. On filesystems that support range reflinks,
>>>> Cassandra can
>>>> >> >> >> also share the underlying extents meaning no new data written.
>>>> Other
>>>> >> >> >> filesystems, including ext4, would copy the already-compressed
>>>> bytes and
>>>> >> >> >> still avoid the row rewrite.
>>>> >> >> >>
>>>> >> >> >> The proposal is staged. It starts with an opt-in `sstablesplit
>>>> >> >> >> --zero-copy` mode for BIG-format SSTables in Cassandra
>>>> 7.0/trunk. Later
>>>> >> >> >> phases add BTI support, secondary indexes, anticompaction, and
>>>> >> >> >> partial-range streaming. Existing implementations remain the
>>>> default and
>>>> >> >> >> provide the fallback for unsupported inputs.
>>>> >> >> >>
>>>> >> >> >> I'd particularly appreciate feedback on:
>>>> >> >> >>
>>>> >> >> >> - The retained-prefix representation and proposed Cassandra
>>>> 7.0 SSTable
>>>> >> >> >> format change
>>>> >> >> >> - Rebuilding or conservatively deriving child metadata without
>>>> decoding
>>>> >> >> >> rows
>>>> >> >> >> - The integrity and performance tradeoff around `Digest.crc32`
>>>> generation
>>>> >> >> >> - The staged rollout, compatibility rules, and fallback
>>>> behavior
>>>> >> >> >> - Any correctness, operational, or filesystem concerns the
>>>> proposal has
>>>> >> >> >> missed
>>>> >> >> >>
>>>> >> >> >> Thanks, and I look forward to the discussion.
>>>> >> >> >>
>>>> >> >> >> Regards,
>>>> >> >> >> Chris Lohfink
>>>> >> >> >>
>>>> >> >> >
>>>> >> >>
>>>> >>
>>>>
>>>
>>
>> --
>> Dmitry Konstantinov
>>
>

-- 
Dmitry Konstantinov

Reply via email to