Hi Benedict, I think the filesystem premise needs correcting. CEP-66 does not link inodes. A range reflink gives each child its own inode while sharing selected physical extents. If the filesystem does not support that operation, the splitter copies the existing compressed chunks instead. That path still avoids decoding and rewriting rows, including on ext4. Reflinking is an optional acceleration, not a requirement for using the feature. I would not claim it makes a completely full disk usable; child components and filesystem metadata still need space. Copying is already the fallback, and the benchmark explains why I would keep the reflink path.
In the reported cold-cache split of a 63 GiB SSTable eight ways, the existing rewrite took 258 seconds and wrote 62.64 GiB. Compressed-byte copying took 259 seconds and wrote 54.85 GiB, although it roughly halved CPU time and greatly reduced heap churn. Reflinking with digest generation enabled took 123 seconds and wrote 57 MiB. Disabling digest generation reduced it to 0.77 seconds, with the verification tradeoff described in the CEP. This is one benchmark, not a universal performance claim, but it is a reason to test the filesystem path rather than discard it on an assumption about typical bandwidth use. Lower write amplification and temporary space pressure matter well before a disk is full. The substantial complexity is constructing valid child SSTables from compression chunks: rebasing indexes, rebuilding derived components, and representing a retained prefix. Copying needs that work too. Both paths produce the same logical child layout, so removing reflinks would not remove the format compatibility work. Phase 1 is an explicit, opt-in offline sstablesplit --zero-copy mode for BIG; online anticompaction is a later phase, initially disabled by default. I am happy to label the opt-in path experimental and make the criteria for broader use explicit. CEP-57 and mutation tracking are valuable efforts, but a future format does not retire BIG and BTI files already in use. Nor does a possible future reduction in anticompaction remove the offline splitting and range-streaming use cases. I do not think those possibilities justify deferring an improvement to current formats. I would welcome evaluating or reusing DataStax’s work, as I said to Branimir. The linked commit adds partial-SSTable reader support with slice metadata; its BigFormat.Version.hasZeroCopyMetadata() returns false. It is therefore not an existing implementation of this first phase for BIG. Workload results and upstreamable code for the same operation would be useful evidence. Removing anticompaction also deserves discussion, but it is a broader design. Repair state and repaired/unrepaired compaction groups are SSTable-wide today. Per-range metadata alone does not explain how partial invalidation and replacement remain correct across reads, compaction, and failures; rewriting the remaining ranges still moves data. I would welcome a concrete proposal, but I cannot treat it as a simpler replacement without that design and measurements. Which correctness or operational risk do you see arising specifically from reflinking, beyond the split representation both paths share? That would let us weigh a concrete cost against the measured benefit. I think you would agree that complexity alone is not a reason to reject a design. You described Accord as “likely the most complex thing we have ever merged to the project,” yet argued that keeping it off by default limited deployment risk. CEP-66 likewise starts as an opt-in offline tool, with any later online use disabled by default. Chris >
