Hi Benedict,

I think the filesystem premise needs correcting. CEP-66 does not link
inodes. A range reflink gives each child its own inode while sharing
selected physical extents. If the filesystem does not support that
operation, the splitter copies the existing compressed chunks instead. That
path still avoids decoding and rewriting rows, including on ext4.
Reflinking is an optional acceleration, not a requirement for using the
feature. I would not claim it makes a completely full disk usable; child
components and filesystem metadata still need space. Copying is already the
fallback, and the benchmark explains why I would keep the reflink path.

In the reported cold-cache split of a 63 GiB SSTable eight ways, the
existing rewrite took 258 seconds and wrote 62.64 GiB. Compressed-byte
copying took 259 seconds and wrote 54.85 GiB, although it roughly halved
CPU time and greatly reduced heap churn. Reflinking with digest generation
enabled took 123 seconds and wrote 57 MiB. Disabling digest generation
reduced it to 0.77 seconds, with the verification tradeoff described in the
CEP. This is one benchmark, not a universal performance claim, but it is a
reason to test the filesystem path rather than discard it on an assumption
about typical bandwidth use. Lower write amplification and temporary space
pressure matter well before a disk is full.

The substantial complexity is constructing valid child SSTables from
compression chunks: rebasing indexes, rebuilding derived components, and
representing a retained prefix. Copying needs that work too. Both paths
produce the same logical child layout, so removing reflinks would not
remove the format compatibility work. Phase 1 is an explicit, opt-in
offline sstablesplit --zero-copy mode for BIG; online anticompaction is a
later phase, initially disabled by default. I am happy to label the opt-in
path experimental and make the criteria for broader use explicit.

CEP-57 and mutation tracking are valuable efforts, but a future format does
not retire BIG and BTI files already in use. Nor does a possible future
reduction in anticompaction remove the offline splitting and
range-streaming use cases. I do not think those possibilities justify
deferring an improvement to current formats. I would welcome evaluating or
reusing DataStax’s work, as I said to Branimir. The linked commit adds
partial-SSTable reader support with slice metadata;
its BigFormat.Version.hasZeroCopyMetadata() returns false. It is therefore
not an existing implementation of this first phase for BIG. Workload
results and upstreamable code for the same operation would be useful
evidence. Removing anticompaction also deserves discussion, but it is a
broader design.

Repair state and repaired/unrepaired compaction groups are SSTable-wide
today. Per-range metadata alone does not explain how partial invalidation
and replacement remain correct across reads, compaction, and failures;
rewriting the remaining ranges still moves data. I would welcome a concrete
proposal, but I cannot treat it as a simpler replacement without that
design and measurements. Which correctness or operational risk do you see
arising specifically from reflinking, beyond the split representation both
paths share? That would let us weigh a concrete cost against the measured
benefit.

I think you would agree that complexity alone is not a reason to reject a
design. You described Accord as “likely the most complex thing we have ever
merged to the project,” yet argued that keeping it off by default limited
deployment risk. CEP-66 likewise starts as an opt-in offline tool, with any
later online use disabled by default.

Chris

>

Reply via email to