Hi Chris, > I think you would agree that complexity alone is not a reason to reject a > design. You described Accord as “likely the most complex thing we have ever > merged to the project,”
Yes, we are discussing cost:benefit. Many proposals are subject to debate, and I am sure you will recall that Accord was by no means exempted from this. This feature primarily benefits a narrow and not recommended use case (EBS), and our roadmap expects to obsolete it. The first draft increases the core codebase size by around 1%, and more for later phases - this isn’t small. > I would welcome a concrete proposal, but I cannot treat it as a simpler > replacement without that design and measurements Fortunately we now have a magic button to press for prototypes. One alternative I have explored that seems compelling to me is simply hard linking the files, updating the first/last key: + Zero CPU/disk cost + Retains digest safety at no CPU cost + Works on all supported systems + Works for all sstable formats + Around 700 LOC total - Space may be reclaimed a little slower, but can configurably encourage/force it on a reasonable schedule This seems like it delivers much of the benefit of phases 1, 2 and 4 for a fraction of the budget? Phase 3 and 5 don’t obviously have a hard dependency on the splitting approach. On 2026/10/01 23:50:15 Chris Lohfink wrote: > Hi Benedict, > > I think the filesystem premise needs correcting. CEP-66 does not link > inodes. A range reflink gives each child its own inode while sharing > selected physical extents. If the filesystem does not support that > operation, the splitter copies the existing compressed chunks instead. That > path still avoids decoding and rewriting rows, including on ext4. > Reflinking is an optional acceleration, not a requirement for using the > feature. I would not claim it makes a completely full disk usable; child > components and filesystem metadata still need space. Copying is already the > fallback, and the benchmark explains why I would keep the reflink path. > > In the reported cold-cache split of a 63 GiB SSTable eight ways, the > existing rewrite took 258 seconds and wrote 62.64 GiB. Compressed-byte > copying took 259 seconds and wrote 54.85 GiB, although it roughly halved > CPU time and greatly reduced heap churn. Reflinking with digest generation > enabled took 123 seconds and wrote 57 MiB. Disabling digest generation > reduced it to 0.77 seconds, with the verification tradeoff described in the > CEP. This is one benchmark, not a universal performance claim, but it is a > reason to test the filesystem path rather than discard it on an assumption > about typical bandwidth use. Lower write amplification and temporary space > pressure matter well before a disk is full. > > The substantial complexity is constructing valid child SSTables from > compression chunks: rebasing indexes, rebuilding derived components, and > representing a retained prefix. Copying needs that work too. Both paths > produce the same logical child layout, so removing reflinks would not > remove the format compatibility work. Phase 1 is an explicit, opt-in > offline sstablesplit --zero-copy mode for BIG; online anticompaction is a > later phase, initially disabled by default. I am happy to label the opt-in > path experimental and make the criteria for broader use explicit. > > CEP-57 and mutation tracking are valuable efforts, but a future format does > not retire BIG and BTI files already in use. Nor does a possible future > reduction in anticompaction remove the offline splitting and > range-streaming use cases. I do not think those possibilities justify > deferring an improvement to current formats. I would welcome evaluating or > reusing DataStax’s work, as I said to Branimir. The linked commit adds > partial-SSTable reader support with slice metadata; > its BigFormat.Version.hasZeroCopyMetadata() returns false. It is therefore > not an existing implementation of this first phase for BIG. Workload > results and upstreamable code for the same operation would be useful > evidence. Removing anticompaction also deserves discussion, but it is a > broader design. > > Repair state and repaired/unrepaired compaction groups are SSTable-wide > today. Per-range metadata alone does not explain how partial invalidation > and replacement remain correct across reads, compaction, and failures; > rewriting the remaining ranges still moves data. I would welcome a concrete > proposal, but I cannot treat it as a simpler replacement without that > design and measurements. Which correctness or operational risk do you see > arising specifically from reflinking, beyond the split representation both > paths share? That would let us weigh a concrete cost against the measured > benefit. > > I think you would agree that complexity alone is not a reason to reject a > design. You described Accord as “likely the most complex thing we have ever > merged to the project,” yet argued that keeping it off by default limited > deployment risk. CEP-66 likewise starts as an opt-in offline tool, with any > later online use disabled by default. > > Chris > > > >
