> Space may be reclaimed a little slower...

This is probably why Chris went with relinking over hard linking. That’s 
potentially a lot of disk space being held onto until compaction gets to it. 
Consider running cleanup after a cluster expansion, you’d have to wait until 
all the of sstables split by cleanup are rewritten before they release a bunch 
of data that's no longer being used.

Regarding complexity, the code needed to support reflinking looks to be pretty 
well isolated, only reachable if you’re using zero copy splitting. On the other 
hand:

> ....but can configurably encourage/force it on a reasonable schedule

moves that complexity into the domains of compaction, configuration, and 
operations. It seems like letting the OS just handle this would be a clear win 
for complexity and operability.

On Tue, Oct 6, 2026, at 4:06 PM, Benedict Elliott Smith wrote:
> >A targeted change that mitigates a large portion of the pain of a feature
> 
> I get the impression my emails are being answered without being read. I 
> proposed an alternative approach that delivers the main benefits to more use 
> cases with less additional logic to maintain.
> 
> On 2026/10/06 22:38:11 Josh McKenzie wrote:
> > > trusting and migrating to new features (especially around consistency) 
> > > will take time, and teams will want to continue using incremental repair 
> > > in the interim.  The CEP providing an opt-in flag for enabling this on 
> > > anticompaction is a good way to prove out and validate the feature.  
> > I second this Andy. Non-trivial features that have consistency implications 
> > historically have had a really long ironing out period for bugs in this 
> > ecosystem. A targeted change that mitigates a large portion of the pain of 
> > a feature that's been a source of serious operational pain for... over a 
> > decade? is non-controversial to me, even if we expect in 1-2 years to have 
> > everyone on an entirely new implementation. That's .5-1.5 years of that 
> > pain being mitigated on the conservative side, assuming things went slowly 
> > on this CEP implementation, best-case on Mutation Tracking, and best-case 
> > on release and adoption.
> > 
> > I think we need to get more comfortable with allowing multiple versions of 
> > things to co-exist (and actually starting to use our experimental flag in 
> > earnest) in this project if we want to make controlled progress on things. 
> > Big swings in software have high risk/reward, and hedging our bets by 
> > improving our existing implementations is just good strategy.
> > 
> > On Tue, Oct 6, 2026, at 5:51 PM, Andy Tolbert wrote:
> > > 
> > > > Indeed. It is exceptionally hard to run k8s without ebs. Possible but 
> > > > requires doing a lot of very unnatural things.
> > > 
> > > Agreed, using Kubernetes without disaggregated storage is incredibly 
> > > difficult and using local disks violates a lot of the value of Kubernetes.
> > > 
> > > I want to make sure that it doesn't get lost that anticompaction is an 
> > > incredibly real problem even with fast NVMEs that causes operational pain 
> > > daily, which would be fantastic to address.
> > > 
> > > If new features (Mutation Tracking) mitigate the need for anticompaction, 
> > > that is great, but trusting and migrating to new features (especially 
> > > around consistency) will take time, and teams will want to continue using 
> > > incremental repair in the interim.  The CEP providing an opt-in flag for 
> > > enabling this on anticompaction is a good way to prove out and validate 
> > > the feature.  Aside from that the CEP also offers value in improving the 
> > > partial-range sstable streaming path as well. 
> > > 
> > > Thanks,
> > > Andy
> > > 
> > > 
> > > On Tue, Oct 6, 2026, at 3:04 PM, Benedict Elliott Smith wrote:
> > > > Not recommended means we do not recommend it, and do not provide 
> > > > guidance for deploying on it. If you believe we *should* recommend EBS 
> > > > we can have that discussion. My view is that a lot of decisions in the 
> > > > database make little sense in this deployment, and should be revisited 
> > > > to support it as a first class deployment option - and somebody needs 
> > > > to do that work. At minimum we need dedicated guidance before 
> > > > considering it one of our supported deployments.
> > > >
> > > > Either way, while this was the motivation, the alternative option’s 
> > > > downside is not really a problem on EBS, so it actually only 
> > > > strengthens the case for pursuing it.
> > > >
> > > >
> > > > On 2026/10/06 19:46:42 Jeff Jirsa wrote:
> > > >> Indeed. It is exceptionally hard to run k8s without ebs. Possible but 
> > > >> requires
> > > >> doing a lot of very unnatural things.
> > > >> 
> > > >>   
> > > >> 
> > > >> Most of the planet is going towards containerized deployments for 
> > > >> relatively
> > > >> obvious reasons.
> > > >> 
> > > >>   
> > > >> 
> > > >> It’s not something we should actively discourage.
> > > >> 
> > > >>   
> > > >> 
> > > >> > On Oct 6, 2026, at 12:32 PM, Jon Haddad <[email protected]> 
> > > >> > wrote:  
> > > >> >  
> > > >> >
> > > >> 
> > > >> > 
> > > >> >
> > > >> > Agreed. Most teams I’ve worked with are on EBS. It’s significantly 
> > > >> > cheaper
> > > >> > for denser storage. I’m running 15-20TB nodes. Ever since we merged
> > > >> > CASSANDRA-15452, its worked really well.
> > > >> >
> > > >> >  
> > > >> >
> > > >> >
> > > >> > Anyone using k8s is probably using EBS.
> > > >> >
> > > >> >  
> > > >> >
> > > >> >
> > > >> > On Tue, Oct 6, 2026 at 12:20 PM Jeff Jirsa
> > > >> > <[[email protected]](mailto:[email protected])> wrote:  
> > > >> >
> > > >> >
> > > >> 
> > > >> >>  
> > > >> >  
> > > >> >  > On Oct 6, 2026, at 4:49 AM, Benedict
> > > >> > <[[email protected]](mailto:[email protected])> wrote:  
> > > >> >  >  
> > > >> >  >   
> > > >> >  >>  
> > > >> >  >> This feature primarily benefits a narrow and not recommended use 
> > > >> > case
> > > >> > (EBS), and our roadmap expects to obsolete it. The first draft 
> > > >> > increases the
> > > >> > core codebase size by around 1%, and more for later phases - this 
> > > >> > isn’t
> > > >> > small.  
> > > >> >  >  
> > > >> >  > Sorry, in my edits for brevity I lost the point I was making 
> > > >> > here. Given
> > > >> > the narrower benefit and lifetime of the feature, and non-trivial 
> > > >> > cost, we
> > > >> > should try to see if we have alternatives that yield a stronger 
> > > >> > cost:benefit
> > > >> > ratio.  
> > > >> >  >  
> > > >> >  > FWIW, I think cheap partial range streaming is an independently 
> > > >> > important
> > > >> > thing we want to achieve, that is of long term benefit to the 
> > > >> > project - and
> > > >> > I am not opposed to these other improvements in principle either.  
> > > >> >  >  
> > > >> >  >> On 6 Oct 2026, at 11:14, Benedict Elliott Smith
> > > >> > <[[email protected]](mailto:[email protected])> wrote:  
> > > >> >  >>  
> > > >> >  >> Hi Chris,  
> > > >> >  >>  
> > > >> >  >>> I think you would agree that complexity alone is not a reason 
> > > >> > to reject
> > > >> > a design. You described Accord as “likely the most complex thing we 
> > > >> > have
> > > >> > ever merged to the project,”  
> > > >> >  >>  
> > > >> >  >> Yes, we are discussing cost:benefit. Many proposals are subject 
> > > >> > to
> > > >> > debate, and I am sure you will recall that Accord was by no means 
> > > >> > exempted
> > > >> > from this.  
> > > >> >  >>  
> > > >> >  >> This feature primarily benefits a narrow and not recommended use 
> > > >> > case
> > > >> > (EBS), and our roadmap expects to obsolete it. The first draft 
> > > >> > increases the
> > > >> > core codebase size by around 1%, and more for later phases - this 
> > > >> > isn’t
> > > >> > small.  
> > > >> >  >>>  
> > > >> >  
> > > >> >  It’s not that narrow, and I dont know that I’d agree about it being 
> > > >> > not-
> > > >> > recommended (I personally put more than one company on it 
> > > >> > intentionally
> > > >> > despite the disk latency because of the cost and operational 
> > > >> > simplicity)  
> > > >> >  
> > > >> >  I’d bet most deployments of Cassandra are on disaggregated disks of 
> > > >> > some
> > > >> > kind (ebs, ceph, azure premium disks, etc).  
> > > >> >  
> > > >> >  Most local nvme offerings are ephemeral in cloud environments and 
> > > >> > much
> > > >> > harder to deal with operationally.
> > > >> 
> > > >>
> > > 
> 

Reply via email to