> This feature primarily benefits a narrow and not recommended use case (EBS), > and our roadmap expects to obsolete it. The first draft increases the core > codebase size by around 1%, and more for later phases - this isn’t small.
Sorry, in my edits for brevity I lost the point I was making here. Given the narrower benefit and lifetime of the feature, and non-trivial cost, we should try to see if we have alternatives that yield a stronger cost:benefit ratio. FWIW, I think cheap partial range streaming is an independently important thing we want to achieve, that is of long term benefit to the project - and I am not opposed to these other improvements in principle either. > On 6 Oct 2026, at 11:14, Benedict Elliott Smith <[email protected]> wrote: > > Hi Chris, > >> I think you would agree that complexity alone is not a reason to reject a >> design. You described Accord as “likely the most complex thing we have ever >> merged to the project,” > > Yes, we are discussing cost:benefit. Many proposals are subject to debate, > and I am sure you will recall that Accord was by no means exempted from this. > > This feature primarily benefits a narrow and not recommended use case (EBS), > and our roadmap expects to obsolete it. The first draft increases the core > codebase size by around 1%, and more for later phases - this isn’t small. > >> I would welcome a concrete proposal, but I cannot treat it as a simpler >> replacement without that design and measurements > > Fortunately we now have a magic button to press for prototypes. One > alternative I have explored that seems compelling to me is simply hard > linking the files, updating the first/last key: > > + Zero CPU/disk cost > + Retains digest safety at no CPU cost > + Works on all supported systems > + Works for all sstable formats > + Around 700 LOC total > - Space may be reclaimed a little slower, but can configurably > encourage/force it on a reasonable schedule > > This seems like it delivers much of the benefit of phases 1, 2 and 4 for a > fraction of the budget? Phase 3 and 5 don’t obviously have a hard dependency > on the splitting approach. > > > >> On 2026/10/01 23:50:15 Chris Lohfink wrote: >> Hi Benedict, >> >> I think the filesystem premise needs correcting. CEP-66 does not link >> inodes. A range reflink gives each child its own inode while sharing >> selected physical extents. If the filesystem does not support that >> operation, the splitter copies the existing compressed chunks instead. That >> path still avoids decoding and rewriting rows, including on ext4. >> Reflinking is an optional acceleration, not a requirement for using the >> feature. I would not claim it makes a completely full disk usable; child >> components and filesystem metadata still need space. Copying is already the >> fallback, and the benchmark explains why I would keep the reflink path. >> >> In the reported cold-cache split of a 63 GiB SSTable eight ways, the >> existing rewrite took 258 seconds and wrote 62.64 GiB. Compressed-byte >> copying took 259 seconds and wrote 54.85 GiB, although it roughly halved >> CPU time and greatly reduced heap churn. Reflinking with digest generation >> enabled took 123 seconds and wrote 57 MiB. Disabling digest generation >> reduced it to 0.77 seconds, with the verification tradeoff described in the >> CEP. This is one benchmark, not a universal performance claim, but it is a >> reason to test the filesystem path rather than discard it on an assumption >> about typical bandwidth use. Lower write amplification and temporary space >> pressure matter well before a disk is full. >> >> The substantial complexity is constructing valid child SSTables from >> compression chunks: rebasing indexes, rebuilding derived components, and >> representing a retained prefix. Copying needs that work too. Both paths >> produce the same logical child layout, so removing reflinks would not >> remove the format compatibility work. Phase 1 is an explicit, opt-in >> offline sstablesplit --zero-copy mode for BIG; online anticompaction is a >> later phase, initially disabled by default. I am happy to label the opt-in >> path experimental and make the criteria for broader use explicit. >> >> CEP-57 and mutation tracking are valuable efforts, but a future format does >> not retire BIG and BTI files already in use. Nor does a possible future >> reduction in anticompaction remove the offline splitting and >> range-streaming use cases. I do not think those possibilities justify >> deferring an improvement to current formats. I would welcome evaluating or >> reusing DataStax’s work, as I said to Branimir. The linked commit adds >> partial-SSTable reader support with slice metadata; >> its BigFormat.Version.hasZeroCopyMetadata() returns false. It is therefore >> not an existing implementation of this first phase for BIG. Workload >> results and upstreamable code for the same operation would be useful >> evidence. Removing anticompaction also deserves discussion, but it is a >> broader design. >> >> Repair state and repaired/unrepaired compaction groups are SSTable-wide >> today. Per-range metadata alone does not explain how partial invalidation >> and replacement remain correct across reads, compaction, and failures; >> rewriting the remaining ranges still moves data. I would welcome a concrete >> proposal, but I cannot treat it as a simpler replacement without that >> design and measurements. Which correctness or operational risk do you see >> arising specifically from reflinking, beyond the split representation both >> paths share? That would let us weigh a concrete cost against the measured >> benefit. >> >> I think you would agree that complexity alone is not a reason to reject a >> design. You described Accord as “likely the most complex thing we have ever >> merged to the project,” yet argued that keeping it off by default limited >> deployment risk. CEP-66 likewise starts as an opt-in offline tool, with any >> later online use disabled by default. >> >> Chris >> >>> >>
