> This feature primarily benefits a narrow and not recommended use case (EBS), 
> and our roadmap expects to obsolete it. The first draft increases the core 
> codebase size by around 1%, and more for later phases - this isn’t small.

Sorry, in my edits for brevity I lost the point I was making here. Given the 
narrower benefit and lifetime of the feature, and non-trivial cost, we should 
try to see if we have alternatives that yield a stronger cost:benefit ratio.

FWIW, I think cheap partial range streaming is an independently important thing 
we want to achieve, that is of long term benefit to the project - and I am not 
opposed to these other improvements in principle either.

> On 6 Oct 2026, at 11:14, Benedict Elliott Smith <[email protected]> wrote:
> 
> Hi Chris,
> 
>> I think you would agree that complexity alone is not a reason to reject a 
>> design. You described Accord as “likely the most complex thing we have ever 
>> merged to the project,”
> 
> Yes, we are discussing cost:benefit. Many proposals are subject to debate, 
> and I am sure you will recall that Accord was by no means exempted from this.
> 
> This feature primarily benefits a narrow and not recommended use case (EBS), 
> and our roadmap expects to obsolete it. The first draft increases the core 
> codebase size by around 1%, and more for later phases - this isn’t small.
> 
>> I would welcome a concrete proposal, but I cannot treat it as a simpler 
>> replacement without that design and measurements
> 
> Fortunately we now have a magic button to press for prototypes. One 
> alternative I have explored that seems compelling to me is simply hard 
> linking the files, updating the first/last key:
> 
> + Zero CPU/disk cost
> + Retains digest safety at no CPU cost
> + Works on all supported systems
> + Works for all sstable formats
> + Around 700 LOC total
> - Space may be reclaimed a little slower, but can configurably 
> encourage/force it on a reasonable schedule
> 
> This seems like it delivers much of the benefit of phases 1, 2 and 4 for a 
> fraction of the budget? Phase 3 and 5 don’t obviously have a hard dependency 
> on the splitting approach.
> 
> 
> 
>> On 2026/10/01 23:50:15 Chris Lohfink wrote:
>> Hi Benedict,
>> 
>> I think the filesystem premise needs correcting. CEP-66 does not link
>> inodes. A range reflink gives each child its own inode while sharing
>> selected physical extents. If the filesystem does not support that
>> operation, the splitter copies the existing compressed chunks instead. That
>> path still avoids decoding and rewriting rows, including on ext4.
>> Reflinking is an optional acceleration, not a requirement for using the
>> feature. I would not claim it makes a completely full disk usable; child
>> components and filesystem metadata still need space. Copying is already the
>> fallback, and the benchmark explains why I would keep the reflink path.
>> 
>> In the reported cold-cache split of a 63 GiB SSTable eight ways, the
>> existing rewrite took 258 seconds and wrote 62.64 GiB. Compressed-byte
>> copying took 259 seconds and wrote 54.85 GiB, although it roughly halved
>> CPU time and greatly reduced heap churn. Reflinking with digest generation
>> enabled took 123 seconds and wrote 57 MiB. Disabling digest generation
>> reduced it to 0.77 seconds, with the verification tradeoff described in the
>> CEP. This is one benchmark, not a universal performance claim, but it is a
>> reason to test the filesystem path rather than discard it on an assumption
>> about typical bandwidth use. Lower write amplification and temporary space
>> pressure matter well before a disk is full.
>> 
>> The substantial complexity is constructing valid child SSTables from
>> compression chunks: rebasing indexes, rebuilding derived components, and
>> representing a retained prefix. Copying needs that work too. Both paths
>> produce the same logical child layout, so removing reflinks would not
>> remove the format compatibility work. Phase 1 is an explicit, opt-in
>> offline sstablesplit --zero-copy mode for BIG; online anticompaction is a
>> later phase, initially disabled by default. I am happy to label the opt-in
>> path experimental and make the criteria for broader use explicit.
>> 
>> CEP-57 and mutation tracking are valuable efforts, but a future format does
>> not retire BIG and BTI files already in use. Nor does a possible future
>> reduction in anticompaction remove the offline splitting and
>> range-streaming use cases. I do not think those possibilities justify
>> deferring an improvement to current formats. I would welcome evaluating or
>> reusing DataStax’s work, as I said to Branimir. The linked commit adds
>> partial-SSTable reader support with slice metadata;
>> its BigFormat.Version.hasZeroCopyMetadata() returns false. It is therefore
>> not an existing implementation of this first phase for BIG. Workload
>> results and upstreamable code for the same operation would be useful
>> evidence. Removing anticompaction also deserves discussion, but it is a
>> broader design.
>> 
>> Repair state and repaired/unrepaired compaction groups are SSTable-wide
>> today. Per-range metadata alone does not explain how partial invalidation
>> and replacement remain correct across reads, compaction, and failures;
>> rewriting the remaining ranges still moves data. I would welcome a concrete
>> proposal, but I cannot treat it as a simpler replacement without that
>> design and measurements. Which correctness or operational risk do you see
>> arising specifically from reflinking, beyond the split representation both
>> paths share? That would let us weigh a concrete cost against the measured
>> benefit.
>> 
>> I think you would agree that complexity alone is not a reason to reject a
>> design. You described Accord as “likely the most complex thing we have ever
>> merged to the project,” yet argued that keeping it off by default limited
>> deployment risk. CEP-66 likewise starts as an opt-in offline tool, with any
>> later online use disabled by default.
>> 
>> Chris
>> 
>>> 
>> 

Reply via email to