Hi Chris,

> I think you would agree that complexity alone is not a reason to reject a 
> design. You described Accord as “likely the most complex thing we have ever 
> merged to the project,” 

Yes, we are discussing cost:benefit. Many proposals are subject to debate, and 
I am sure you will recall that Accord was by no means exempted from this.

This feature primarily benefits a narrow and not recommended use case (EBS), 
and our roadmap expects to obsolete it. The first draft increases the core 
codebase size by around 1%, and more for later phases - this isn’t small.

> I would welcome a concrete proposal, but I cannot treat it as a simpler 
> replacement without that design and measurements

Fortunately we now have a magic button to press for prototypes. One alternative 
I have explored that seems compelling to me is simply hard linking the files, 
updating the first/last key:

 + Zero CPU/disk cost
 + Retains digest safety at no CPU cost
 + Works on all supported systems
 + Works for all sstable formats
 + Around 700 LOC total
 - Space may be reclaimed a little slower, but can configurably encourage/force 
it on a reasonable schedule

This seems like it delivers much of the benefit of phases 1, 2 and 4 for a 
fraction of the budget? Phase 3 and 5 don’t obviously have a hard dependency on 
the splitting approach.



On 2026/10/01 23:50:15 Chris Lohfink wrote:
> Hi Benedict,
> 
> I think the filesystem premise needs correcting. CEP-66 does not link
> inodes. A range reflink gives each child its own inode while sharing
> selected physical extents. If the filesystem does not support that
> operation, the splitter copies the existing compressed chunks instead. That
> path still avoids decoding and rewriting rows, including on ext4.
> Reflinking is an optional acceleration, not a requirement for using the
> feature. I would not claim it makes a completely full disk usable; child
> components and filesystem metadata still need space. Copying is already the
> fallback, and the benchmark explains why I would keep the reflink path.
> 
> In the reported cold-cache split of a 63 GiB SSTable eight ways, the
> existing rewrite took 258 seconds and wrote 62.64 GiB. Compressed-byte
> copying took 259 seconds and wrote 54.85 GiB, although it roughly halved
> CPU time and greatly reduced heap churn. Reflinking with digest generation
> enabled took 123 seconds and wrote 57 MiB. Disabling digest generation
> reduced it to 0.77 seconds, with the verification tradeoff described in the
> CEP. This is one benchmark, not a universal performance claim, but it is a
> reason to test the filesystem path rather than discard it on an assumption
> about typical bandwidth use. Lower write amplification and temporary space
> pressure matter well before a disk is full.
> 
> The substantial complexity is constructing valid child SSTables from
> compression chunks: rebasing indexes, rebuilding derived components, and
> representing a retained prefix. Copying needs that work too. Both paths
> produce the same logical child layout, so removing reflinks would not
> remove the format compatibility work. Phase 1 is an explicit, opt-in
> offline sstablesplit --zero-copy mode for BIG; online anticompaction is a
> later phase, initially disabled by default. I am happy to label the opt-in
> path experimental and make the criteria for broader use explicit.
> 
> CEP-57 and mutation tracking are valuable efforts, but a future format does
> not retire BIG and BTI files already in use. Nor does a possible future
> reduction in anticompaction remove the offline splitting and
> range-streaming use cases. I do not think those possibilities justify
> deferring an improvement to current formats. I would welcome evaluating or
> reusing DataStax’s work, as I said to Branimir. The linked commit adds
> partial-SSTable reader support with slice metadata;
> its BigFormat.Version.hasZeroCopyMetadata() returns false. It is therefore
> not an existing implementation of this first phase for BIG. Workload
> results and upstreamable code for the same operation would be useful
> evidence. Removing anticompaction also deserves discussion, but it is a
> broader design.
> 
> Repair state and repaired/unrepaired compaction groups are SSTable-wide
> today. Per-range metadata alone does not explain how partial invalidation
> and replacement remain correct across reads, compaction, and failures;
> rewriting the remaining ranges still moves data. I would welcome a concrete
> proposal, but I cannot treat it as a simpler replacement without that
> design and measurements. Which correctness or operational risk do you see
> arising specifically from reflinking, beyond the split representation both
> paths share? That would let us weigh a concrete cost against the measured
> benefit.
> 
> I think you would agree that complexity alone is not a reason to reject a
> design. You described Accord as “likely the most complex thing we have ever
> merged to the project,” yet argued that keeping it off by default limited
> deployment risk. CEP-66 likewise starts as an opt-in offline tool, with any
> later online use disabled by default.
> 
> Chris
> 
> >
> 

Reply via email to