To my eye, 2 and 3 are redundant. I’m all about KISS here. Patrick
> On Sep 24, 2026, at 5:40 PM, Caleb Rackliffe <[email protected]> wrote: > Blake, I’d support those 3 guidelines. > > One hostile reaction to them would, I’m guessing, be that guideline number 1 > is too weak, and that the burden of quality would continue to rest entirely > on committer reviewers. I suppose my reply there would be that anyone who > continually spams the project with patches they do not fully understand wound > incur a pretty heavy reputational penalty and the problem would more or less > sort itself out. > > >> On Sep 24, 2026, at 6:07 PM, Benedict Elliott Smith <[email protected]> >> wrote: >> >> For the record Josh, I am including pre-3.0. I was involved in customer >> support escalations for 1.2, 2.0 and 2.1: the software was shoddy, to put it >> mildly. 8099 was a symptom, not the disease. >> >> The project had a culture of doing stuff without sufficient care or >> consideration. This is by no means a simple matter of needing more testing >> or less timeline pressure, nor is it easy to define a "bar" to ensure it >> does not regress. If we had simple metrics, we would have settled it a long >> time ago. >> >> You can still see customer scars crop up on forum discussions periodically. >> >> >> On 2026/09/24 19:01:02 Josh McKenzie wrote: >>>> Apache Cassandra was fundamentally undeployable for four years between Nov >>>> 2015 - 2019. >>> CASSANDRA-8099 was a maximal manifestation of a specific approach to >>> engineering and calendar constraints we've seen time and again on the >>> project; I don't want us to conflate things here. That was a herculean >>> monolithic body of work performed in inhuman conditions (in vim!) that was >>> ultimately so invasive, all the unit tests in the code-base were commented >>> out and Jake and I spent a grueling 1.5-2 months hand-rewriting basically >>> all the unit tests in that code-base to get things to even build and run, >>> much less pass. >>> Massive blast radius changes that are un-sustainably complex, under-tested, >>> where we don't property, fuzz, check coverage, check complexity, or A/B >>> compare against a known good system (i.e. pre/post) correctness testing are >>> going to destabilize the database at any time under any regime of tooling. >>> We certainly could speed-run our way back into destabilization with LLM's >>> as they are a force multiplier for both the good and the bad of one's >>> engineering practices, but that's a solvable problem by better defining >>> what our bars of quality are (Definition of Done anyone?) and holding >>> ourselves accountable to delivering at that bar. >>> On Thu, Sep 24, 2026, at 2:40 PM, David Capwell via dev wrote: >>>>> if you really want to pursue it I would ask that we do it offline to >>>>> avoid polluting an already busy conversation >>>> People are directly responding saying that they feel discrimination >>>> currently and that the policy tries to codify that discrimination, so I >>>> feel its 100% on topic. This thread has presented 0 evidence that LLM >>>> usage has lowered the quality of contributions merged and has so far been >>>> vibes and feeling; I have yet to see any evidence to justify such >>>> discrimination so I will keep pushing back until such evidence is >>>> presented so we can have a informed debate. >>>>> namely how we handle shallow and localised bug fixes. I would be happy >>>>> adding a clear entry to the “Permitted” section for this. No doubt there >>>>> are many refinements needed to the Restricted text as well, that might >>>>> also capture some of your concerns. >>>> I will not sign off on cherry picking areas where "safe" to use a tool; so >>>> no it does not capture my concerns. >>>>> I don't know if everyone remembers, but ten years ago Cassandra was full >>>>> of serious correctness and stability issues. Despite developing it, I >>>>> would not have run it myself or recommend that anyone use it. We have dug >>>>> ourselves out of that hole, but it took years of discipline and effort, >>>>> and we're still (deservedly) recovering our reputation. >>>>> Let's use this new technology to improve the quality of our >>>>> contributions, not squander our hard-earned gains in the name of speed. >>>>> It will be hard to recover our reputation a second time. >>>> I do recall the 3.x line and put in a significant amount of effort to >>>> harden it. There were behaviors I noticed after joining Cassandra that I >>>> feel directly contributed to 3.0 and the decade of catch up; behaviors >>>> that still linger in parts of the community today. >>>> As I look on trunk and look at committed code and trace back to PRs and >>>> JIRA I see the following: >>>> • large patches approved without comments >>>> • 0 evidence that tests were run >>>> I then look at our CI and see tests failing for months. As you start to >>>> triage you start to see some of them show real issues; yet they linger for >>>> months not being addressed... when CI is unstable it takes a lot of effort >>>> to triage "did my patch break the test", and I have seen time and time >>>> again people do not put in that effort, and shrug off as "its just a >>>> flakey test"; then our CI failure rate grows. >>>> Non of this has anything to do with LLMs but LLMs running in this >>>> environment is far more dangerous as there are not checks in place to >>>> "hold the bar". I am all for raising the bar universally; expecting both >>>> humans and LLMs to match that bar. >>>>> `I’d support something that boils down to roughly this: >>>>> 1.) 2 committers must understand an LLM-assisted change before it >>>>> commits. (Perhaps separately we can explore the question of why we >>>>> haven’t added any new committers to the core project for about a year. >>>>> I’m also still not entirely sure if it’s acceptable within our guidelines >>>>> for a committer to +1 a patch after delegating review.) >>>>> 2.) Patch authors must demonstrate enough understanding to discuss their >>>>> own patch, whether or not parts of it are generated by an LLM. >>>>> 3.) The “Assisted-by” tag should be used to indicate any non-trivial LLM >>>>> usage in the generation of a patch, just like we have used Co-authored-by >>>>> historically. >>>>> 4.) Comments and other things that aren't the actual code (but could sow >>>>> confusion) should be held to the same standard we'd expect from a human >>>>> writer. If we don’t yet agree on that standard, we can formalize enough >>>>> of it to guide both humans and LLMs. >>> ` >>>> I can get behind this proposal but i would tweak it as 1/2 i don't think >>>> really need to special case LLM usage >>>> 1. 2 committers must understand the change before it commits. >>>> 2. Patch authors must demonstrate enough understanding to discuss their >>>> own patch >>>> Nothing about those 2 need to be scoped to LLM usage and honestly matches >>>> most PMCs I have talked to understanding of our bar (as Benedict pointed >>>> out, the actual wording could be interpreted to allow rubber stamping from >>>> committer) >>>> As for 3 I am cool with this. ASF recommends the same (it says >>>> `Generated-by` but that discount's the human's effort) as its useful for >>>> audits and tooling. Having `Assisted-by` tag should not imply anything >>>> about the committed patch as it should have gone through the same bar we >>>> all expect; its just for tools auditing. >>>>> On Sep 24, 2026, at 10:37 AM, Aleksey Yeshchenko via dev >>>>> <[email protected]> wrote: >>>>> Meant "doing away with", sorry. Non-native speaker with a headache here. >>>>> Thanks Caleb for spotting. >>>>>> On 24 Sep 2026, at 17:35, Aleksey Yeshchenko via dev >>>>>> <[email protected]> wrote: >>>>>> P.S. I assume it's obvious from the text above that I don't believe that >>>>>> getting away with human code review is a viable option. >>>>>> On 24 Sep 2026, at 17:37, Štefan Miklošovič <[email protected]> >>>>>> wrote: >>>>>> Good call on checkerframework, we even have a patch for it. Work of >>>>>> Jacek Lewandowski. We might just drive it to completion. Using AI for >>>>>> finishing it would be quite ironic. >>>>>> (1) https://github.com/apache/cassandra/pull/2370 >>>>>> On Thu, Sep 24, 2026 at 6:19 PM Jon Haddad <[email protected]> >>>>>> wrote: >>>>>>> There are some really good points being brought up about stability of >>>>>>> the codebase, maintainability, quality of reviews, correctness bugs, >>>>>>> and I agree with all of them. I think it would be helpful to take a >>>>>>> step back and consider how those bugs got there in the first place, how >>>>>>> they were fixed, and what we could do to further advance the codebase >>>>>>> so they don't creep back. LLMs can be used either with very tight >>>>>>> guardrails, or in what's effectively YOLO mode, and there's a big >>>>>>> difference in the quality of the results you get. >>>>>>> One thing to keep in mind, a lot of the initial code in C* was added >>>>>>> without comprehensive testing. I hope we can all agree that it's a lot >>>>>>> easier to break code that doesn't have high quality tests. During the >>>>>>> code freeze, a lot of people people spent several years relentlessly >>>>>>> finding and fixing bugs. This was probably a pretty frustrating time >>>>>>> for anyone who was focused on fixing other people's bugs when they >>>>>>> wanted to build features. I think we should recognize the effort here >>>>>>> and appreciate the foundation that the project stands on now. I can >>>>>>> understand how anyone involved with this effort would be apprehensive >>>>>>> about seeing years of their life swept away by an agent that was driven >>>>>>> by goal seeking to remove all the tests that it broke instead of fixing >>>>>>> them. >>>>>>> When I picked up the work to improve cursor compaction, the first thing >>>>>>> I asked myself was how can I make sure I don't break this? How do I >>>>>>> even know it works properly? There were some tricky parts to the code, >>>>>>> and I really didn't want to come in and immediately break stuff. >>>>>>> That's why I started with an entire patch dedicated to adding test >>>>>>> infra to it. 90% of the patch was tests, and in my other cursor >>>>>>> patches, it remains *at least* 80% of my patches. It was a *lot* >>>>>>> faster to add almost 10K lines of tests that handled a byte for byte >>>>>>> differential testing paired with harry to find over 30 bugs that caused >>>>>>> cursor to corrupt results. Range tombstones alone were at least a >>>>>>> dozen bugs, but I also found issues with static columns, reverse >>>>>>> ordering, etc. Randomizing schemas and data in burn tests to generate >>>>>>> different shapes of data, to ensure they all result in the same output >>>>>>> at the end. JMH tests to ensure there weren't performance regressions, >>>>>>> hours of profiling. These were all a *lot* easier to do with the LLM >>>>>>> helping me out. In the process I've found bugs that have been >>>>>>> lingering in the codebase for years. >>>>>>> That's a long story, but hopefully we all agree that having >>>>>>> comprehensive tests is a great way to ensure that both humans and LLMs >>>>>>> don't break things that are working. >>>>>>> The lesson: we need to keep improving our testing. Everything that we >>>>>>> touch, should be left in a better state than how we found it with >>>>>>> regard to test coverage. >>>>>>> Test coverage isn't everything though, there's always little subtle >>>>>>> bugs that don't get found in testing, that can slip in despite our best >>>>>>> efforts. It's debatable if humans will be as good as agents for coding >>>>>>> in the long term, for spotting small defects. I sincerely doubt it. >>>>>>> For the time being though, we still have people involved. It's probably >>>>>>> a good time to start using more static analysis tools to identify >>>>>>> problematic code and to add this to CI. Dmitry had a suggestion >>>>>>> recently for checkerframework to detect leaking contexts, a problem he >>>>>>> spotted when reviewing my branch. It would be great to have that >>>>>>> integrated into our CI and dev workflow so we can simply avoid an >>>>>>> entire class of bugs. >>>>>>> There's also PMD, which is excellent for finding code that can be hard >>>>>>> to understand. I *highly* suggest you all run PMD to analyze for >>>>>>> cognitive complexity and high npath scores. This was made popular by >>>>>>> the folks at Sonar and I've found it to be an excellent feedback >>>>>>> mechanism for structuring code. The default max they set is 15, which >>>>>>> is the point where it starts to become difficult to verify something >>>>>>> works without making a massive investment. We've got areas in the >>>>>>> codebase that are in the hundreds, and some parts even higher. These >>>>>>> have been contributed by humans, and are all high risk points for both >>>>>>> humans and agents to start messing around with. They're also in some >>>>>>> fairly critical areas that are very likely to break, so I understand >>>>>>> why people would not want an agent anywhere near it. >>>>>>> Unfortunately, it's not an easy problem to address. There's so many >>>>>>> places where the code is structured in a way that has so many branches, >>>>>>> so many conditions, that it's effectively impossible for a human to >>>>>>> understand, creating a fear of messing around in it. There's plenty of >>>>>>> areas that deserve extreme scrutiny, and we should be careful of what >>>>>>> we add, whether it's human or agent. >>>>>>> The codebase today requires a high degree of internal knowledge to >>>>>>> navigate. There's land mines everywhere. We should be looking to make >>>>>>> conscious improvements by moving the code forward, so it's easier to >>>>>>> make changes to small, well tested components with minimal side >>>>>>> effects. Not making it harder for people to use the tools that aid in >>>>>>> that process. >>>>>>> Here's what we could do to achieve the underlying goal of not breaking >>>>>>> the DB: >>>>>>> Add cognitive complexlity and npath via PMD as a feedback mechanism. >>>>>>> Code that's hard to understand is hard to review. It's also hard to >>>>>>> test. Let's break down the complex code so more people can contribute, >>>>>>> safely. >>>>>>> Add checkerframework to our tooling, >>>>>>> Properly annotate the codebase for it and reduce the surface area that >>>>>>> things can break. Less brittle codebase = we can move faster. >>>>>>> Use jacoco to find areas of the codebase with poor testing. >>>>>>> Let's improve the test coverage there, LLMs are great for this. We >>>>>>> have a ton of static tests, these can become more dynamic, >>>>>>> parameterized, and leverage harry. >>>>>>> Refactor parts of the codebase that have high cognitive complexlity and >>>>>>> NPath scores. >>>>>>> This should be lowered over time to meet some high watermark, say 25 >>>>>>> maximum, although I'd prefer 15 which is where the Sonar folks settled. >>>>>>> Move forward moving the codebase to a more modular structure >>>>>>> We've talked about Gradle on and off - but it can really be a huge help >>>>>>> with incremental, modular builds. This is pretty easy to do with an >>>>>>> agent and we could have it done in a couple days. >>>>>>> Enforce boundaries with ArchUnit >>>>>>> If we want to enforce certain code boundaries, this is the way to do >>>>>>> it. Should not be part of manual review. >>>>>>> Add LLM review for all incoming PRs before a human >>>>>>> The goal here is to automate the initial part of the review process >>>>>>> that reviewers should spot, and raise the bar for the initial >>>>>>> contribution. When the code gets reviewed by a human, it should >>>>>>> already have passed a large variety of initial checks. This should >>>>>>> shorten the review cycle and result in higher quality patches. I've >>>>>>> had Claude reviewing all my PRs in my personal projects for a while now >>>>>>> and it consistently gives great feedback that I almost always >>>>>>> incorporate. >>>>>>> In my ideal world, we'd also auto-format all code >>>>>>> Consistent formatting throughout the codebase would be amazing, but >>>>>>> that's just one man's dream. >>>>>>> Hopefully there's at least a couple things in this list we could move >>>>>>> forward with in the short term, as it'll help improve the code quality >>>>>>> regardless of how it's created. >>>>>>> Jon >>>>>>> https://checkerframework.org/manual/#aliasing-leaking-contexts >>>>>>> https://www.sonarsource.com/docs/CognitiveComplexity.pdf >>>>>>> https://pmd.github.io/pmd/pmd_rules_java_design.html >>>>>>> On Thu, Sep 24, 2026 at 7:38 AM C. Scott Andreas <[email protected]> >>>>>>> wrote: >>>>>>>> From Benedict: >>>>>>>> “I don't know if everyone remembers, but ten years ago Cassandra was >>>>>>>> full of serious correctness and stability issues. Despite developing >>>>>>>> it, I would not have run it myself or recommend that anyone use it. We >>>>>>>> have dug ourselves out of that hole, but it took years of discipline >>>>>>>> and effort, and we're still (deservedly) recovering our reputation.” >>>>>>>> Expanding on this point for those who may not have been active in the >>>>>>>> project at this time — >>>>>>>> Apache Cassandra was fundamentally undeployable for four years between >>>>>>>> Nov 2015 - 2019. The database literally lost data if you ran a >>>>>>>> read-only SELECT query ordered descending (C-14513, C-14515). If you >>>>>>>> haven’t read these tickets before, please take a moment to do so. >>>>>>>> It took years of careful work via property-based testing, fuzzing, and >>>>>>>> deterministic simulation to restore Cassandra’s status as a usable >>>>>>>> system of record. Once 14513 and 14515 were identified, nearly 30 >>>>>>>> additional critical data loss and incorrect response bugs were >>>>>>>> identified. >>>>>>>> It is essential for the project’s future that we don’t regress to this >>>>>>>> state chasing AI-generated features motivated by fear. The fact that >>>>>>>> examples cited in this thread which boast shiny features but have >>>>>>>> critical shortcomings unknown to their author supports this argument. >>>>>>>> The most common path for large corpuses of AI-generated software is >>>>>>>> elation and reveling in a feature matrix, followed by abandonment. >>>>>>>> I endorse this point: >>>>>>>> “Let's use this new technology to improve the quality of our >>>>>>>> contributions, not squander our hard-earned gains in the name of >>>>>>>> speed. It will be hard to recover our reputation a second time.” >>>>>>>> Patrick, I don’t want your note regarding a TCM issue to go >>>>>>>> unaddressed. Please file a Jira ticket and the patch if you like. I >>>>>>>> can’t comment on the patch as I haven’t seen it, but together we will >>>>>>>> solve the problem. >>>>>>>> – Scott >>>>>>>>> On Sep 24, 2026, at 4:01 AM, Benedict Elliott Smith >>>>>>>>> <[email protected]> wrote: >>>>>>>>> Hi Patrick, >>>>>>>>> As I mentioned in my reply to David, I would be happy to create a >>>>>>>>> carve out for shallow and localised bug fixes in the "Permitted" >>>>>>>>> section. Would this alleviate some of your concerns regarding your >>>>>>>>> ability to contribute to the project? >>>>>>>>> I appreciate your pointing out Ferrosa's Accord implementation >>>>>>>>> however, as it is a *great* example of the problems we're leaping >>>>>>>>> into. I took a look, and within about 30s found that the protocol is >>>>>>>>> fundamentally incorrect, having failed to address CASSANDRA-18365. >>>>>>>>> This is despite claiming to be tested with Jepsen that should in >>>>>>>>> principle find this fault. I followed up by using Claude to >>>>>>>>> interrogate the implementation further, and immediately found other >>>>>>>>> serious correctness issues. >>>>>>>>> I use LLMs daily now to help facilitate Accord development, and while >>>>>>>>> they are powerful they are NOT able to author the code themselves, >>>>>>>>> even when building upon a strong human-authored foundation. >>>>>>>>> I don't know if everyone remembers, but ten years ago Cassandra was >>>>>>>>> full of serious correctness and stability issues. Despite developing >>>>>>>>> it, I would not have run it myself or recommend that anyone use it. >>>>>>>>> We have dug ourselves out of that hole, but it took years of >>>>>>>>> discipline and effort, and we're still (deservedly) recovering our >>>>>>>>> reputation. >>>>>>>>> Let's use this new technology to improve the quality of our >>>>>>>>> contributions, not squander our hard-earned gains in the name of >>>>>>>>> speed. It will be hard to recover our reputation a second time. >>>>>>>>>> On 2026/09/23 19:16:17 Patrick McFadin wrote: >>>>>>>>>> I was waiting for this moment to hit our project and I'm glad we're >>>>>>>>>> here. I >>>>>>>>>> am deeply concerned for our project and its future, as we have >>>>>>>>>> increasingly >>>>>>>>>> made it difficult to contribute. I had hoped that this new era of >>>>>>>>>> software tools powered by AI would expand the project's reach and >>>>>>>>>> bring >>>>>>>>>> more diverse thoughts and ideas. This policy proposal is the exact >>>>>>>>>> opposite >>>>>>>>>> of what we need. We have been sitting on a Cassandra 6 release alpha >>>>>>>>>> for >>>>>>>>>> months. We need to accelerate and embrace new ways of being or be >>>>>>>>>> left >>>>>>>>>> behind. As I read that policy, my first and gut level reactions: >>>>>>>>>> - It comes across as elitist and class protectionism. Committer >>>>>>>>>> should not >>>>>>>>>> be special but this proposal makes that designation even more sacred. >>>>>>>>>> - It signals that our project is so fragile that only a few people >>>>>>>>>> "Really >>>>>>>>>> understand it" That's some SQLite vibes right there. >>>>>>>>>> - Trying to fix a problem that doesn't exist >>>>>>>>>> Sadly, i think this policy change would also exclude a lot of >>>>>>>>>> comitters. >>>>>>>>>> We aren't alone in this moment. The Linux project just went through >>>>>>>>>> this. You can find the thread with a simple Google, but similar hard >>>>>>>>>> feelings were being expressed "AI is going to ruin our project!", >>>>>>>>>> "The >>>>>>>>>> unwashed masses are going to contribute terrible code!", "We have to >>>>>>>>>> protect our precious status as Linux maintainers!" Linus being >>>>>>>>>> Linus, was >>>>>>>>>> deeply invloved and they adopted a super simple statement that >>>>>>>>>> covers all >>>>>>>>>> bases. Human or Human using AI. “You are expected to understand and >>>>>>>>>> to be >>>>>>>>>> able to defend everything you submit.” Love that. >>>>>>>>>> In the larger picture, I'll restate. I'm worried for our project. In >>>>>>>>>> late >>>>>>>>>> 2025(Opus 4.5 IYKYK), early 2026, AI coding LLMs turned a real >>>>>>>>>> corner and >>>>>>>>>> in the hands of somebody that knows how to build software, this tool >>>>>>>>>> is >>>>>>>>>> like jet fuel. Here's some examples of new projects being hyper >>>>>>>>>> fueled by >>>>>>>>>> AI coding tools. >>>>>>>>>> Apache Iggy - Complete rust replacement of kafka. Crazy fast velocity >>>>>>>>>> Turso - Rust re-write of SQLite >>>>>>>>>> Bun - Rust re-write of itself from Zig. >>>>>>>>>> Think this couldn't happen to us? Already has: >>>>>>>>>> https://github.com/ferrosadb/ferrosa. Ben is using it to power his >>>>>>>>>> own >>>>>>>>>> startup, but it was him alone using a ton of local AI coding agents. >>>>>>>>>> He >>>>>>>>>> even implemented Accord. Yeah... >>>>>>>>>> The cracks are already starting to show. There is a black market >>>>>>>>>> economy of >>>>>>>>>> Cassandra patches happening now. Not going to name names or call >>>>>>>>>> people >>>>>>>>>> out, but there are fixes and optimizations living in branches >>>>>>>>>> outside of >>>>>>>>>> the Cassandra project. Why? I'll use myself as an example. I fixed a >>>>>>>>>> nasty >>>>>>>>>> bug I ran into with TCM a few weeks ago. Wrote the tests. It passes >>>>>>>>>> CI and >>>>>>>>>> lives in my personal branch. I'm sitting here really wondering if I >>>>>>>>>> want to >>>>>>>>>> go through the ritual humiliation of being roasted for using AI to >>>>>>>>>> fix it. >>>>>>>>>> Me. I am worried about contrinuting code the Cassandra. What the >>>>>>>>>> hell does >>>>>>>>>> that say? >>>>>>>>>> I have my CQLite project that I've been doing a release around once a >>>>>>>>>> month. I would love to donate that to the Cassandra project but I >>>>>>>>>> wouldn't >>>>>>>>>> if it essentially killed any progress. >>>>>>>>>> My larger counter proposal would be to: >>>>>>>>>> - Adopt the “You are expected to understand and to be able to defend >>>>>>>>>> everything you submit.” approach the Linux project has adopted. >>>>>>>>>> - Loosen up the contributor process and our worry on trunk. Let 1000 >>>>>>>>>> flowers bloom and bring it in. >>>>>>>>>> - And finally, to give some people more peace of mind and open more >>>>>>>>>> doors, >>>>>>>>>> adopt what other projects have done and provide more pluggability. >>>>>>>>>> Let new >>>>>>>>>> ideas have an easy place to connect. >>>>>>>>>> We are at a fork in the road. What are we going to do? And then I >>>>>>>>>> have to >>>>>>>>>> ask myself, what am I going to do as a contributor? >>>>>>>>>> Patrick >>>>>>>>>> On Wed, Sep 23, 2026 at 6:16 AM Blake Eggleston >>>>>>>>>> <[email protected]> >>>>>>>>>> wrote: >>>>>>>>>>> I’m not necessarily opposed to having a policy, but so far we have >>>>>>>>>>> some >>>>>>>>>>> specific proposals addressing a problem statement that’s very >>>>>>>>>>> nebulous. >>>>>>>>>>> What is the community failing to do on its own that we’re trying to >>>>>>>>>>> correct >>>>>>>>>>> with policy? What outcomes are we trying to create or prevent? >>>>>>>>>>> Having some >>>>>>>>>>> examples and specific problems to discuss would help focus the >>>>>>>>>>> conversation. >>>>>>>>>>>> On Wed, Sep 23, 2026, at 4:34 AM, Shailaja Koppu via dev wrote: >>>>>>>>>>> Benedict, >>>>>>>>>>> Thanks for clarifying. My concern still remains. This criteria >>>>>>>>>>> would be >>>>>>>>>>> difficult to define and apply consistently. What counts as >>>>>>>>>>> “similar” scope >>>>>>>>>>> or area, “mostly correct,” or sufficiently independent work? More >>>>>>>>>>> importantly, how do we prevent such vague criteria from creating an >>>>>>>>>>> informal hierarchy where some contributors work is routinely >>>>>>>>>>> accepted while >>>>>>>>>>> others is routinely rejected? >>>>>>>>>>> If the intent is to limit AI-assisted code changes to Cassandra >>>>>>>>>>> contributors, or to contributors who have previously worked in that >>>>>>>>>>> component without AI, that would at least be clear and enforceable. >>>>>>>>>>>> On Sep 23, 2026, at 12:01 PM, Benedict Elliott Smith < >>>>>>>>>>> [email protected]> wrote: >>>>>>>>>>>> Core code changes >>>>>>>>>>>> Chris: Do you object to the first or second line you quote? >>>>>>>>>>>> Because the >>>>>>>>>>> first line is effectively motivation for the second line, and can be >>>>>>>>>>> removed (or more clearly combined). If it’s the second line, then I >>>>>>>>>>> do not >>>>>>>>>>> think this is an unreasonable expectation, and we can get into a >>>>>>>>>>> proper >>>>>>>>>>> debate about it. >>>>>>>>>>>> Shailaja, since you only snipped the first sentence, your concerns >>>>>>>>>>>> might >>>>>>>>>>> also be mostly answered by this clarification? “Minimal third-party >>>>>>>>>>> guidance” implies you have some concerns about the second line, but >>>>>>>>>>> all of >>>>>>>>>>> our policies have some ambiguity because legalese is even worse. I >>>>>>>>>>> don’t >>>>>>>>>>> think the ambiguity here would be challenging to navigate though we >>>>>>>>>>> can >>>>>>>>>>> certainly refine it. This specific snippet is meant to convey an >>>>>>>>>>> expectation that a contributor has autonomously produced patches of >>>>>>>>>>> similar >>>>>>>>>>> scope that were mostly correct, so that they have demonstrated the >>>>>>>>>>> level of >>>>>>>>>>> understanding necessary to guide another party to a successful >>>>>>>>>>> patch (i.e. >>>>>>>>>>> an LLM in this case). >>>>>>>>>>>> On 2026/09/23 10:54:16 Benedict Elliott Smith wrote: >>>>>>>>>>>>> Thanks everyone for your input so far. I’ll respond in brief to >>>>>>>>>>>>> the >>>>>>>>>>> main themes, in (mostly) separate emails so they can each have >>>>>>>>>>> their own >>>>>>>>>>> debate chain. >>>>>>>>>>>>> Should we have a policy (Blake/Josh*/Jon/Dinesh) >>>>>>>>>>>>> I think we would all agree that LLMs represent the biggest change >>>>>>>>>>>>> to >>>>>>>>>>> this community (and software more generally) since its inception, >>>>>>>>>>> and we >>>>>>>>>>> all now have enough experience with the technology to have formed >>>>>>>>>>> opinions >>>>>>>>>>> about how it is best managed. We also evidently have not all >>>>>>>>>>> arrived at the >>>>>>>>>>> same conclusions. In this situation, it would be an abdication of >>>>>>>>>>> our >>>>>>>>>>> responsibilities as a management committee to not agree *some* >>>>>>>>>>> policy. >>>>>>>>>>>>> I intend to conduct straw polls as the discussion evolves, so if >>>>>>>>>>>>> you >>>>>>>>>>> prefer an alternative policy - or modifications to this policy - I >>>>>>>>>>> would >>>>>>>>>>> encourage you to make those alternative proposals. >>>>>>>>>>>>> *Veto/Consensus (Josh) >>>>>>>>>>>>> It was fair to call out my poor use of language on this topic, so >>>>>>>>>>>>> let >>>>>>>>>>> me rephrase a little. The community is built on consensus, and work >>>>>>>>>>> should >>>>>>>>>>> not be merged when there are outstanding concerns to address. The >>>>>>>>>>> explicit >>>>>>>>>>> -1 should only be used rarely, because the prior expectation should >>>>>>>>>>> prevent >>>>>>>>>>> it ever being needed. I (and others) have outstanding concerns on >>>>>>>>>>> LLM >>>>>>>>>>> generated work that can only be addressed through this process >>>>>>>>>>> right here, >>>>>>>>>>> so to merge such work while maintaining the community’s consensus >>>>>>>>>>> we must >>>>>>>>>>> agree some policy. >>>>>>>>>>>>> On 2026/09/23 09:58:27 Shailaja Koppu via dev wrote: >>>>>>>>>>>>>> I am strongly -1 on this >>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by >>>>>>>>>>>>>> contributors >>>>>>>>>>> with demonstrated expertise >>>>>>>>>>>>>> That creates a new, subjective privileged class of contributors >>>>>>>>>>>>>> and >>>>>>>>>>> turns a tool choice into an eligibility test. Who decides whether >>>>>>>>>>> expertise >>>>>>>>>>> has been “demonstrated,” what counts as “minimal third-party >>>>>>>>>>> guidance,” and >>>>>>>>>>> how could those judgments be applied consistently or fairly? >>>>>>>>>>>>>> Apache already has a better model, anyone may contribute, trust >>>>>>>>>>>>>> and >>>>>>>>>>> additional repository privileges are earned transparently over >>>>>>>>>>> time. The >>>>>>>>>>> ASF describes its communities as flat, and says that newcomer ideas >>>>>>>>>>> have as >>>>>>>>>>> much input as those from original creators. We should not add a >>>>>>>>>>> separate, >>>>>>>>>>> informal hierarchy in which certain people may use common >>>>>>>>>>> development tools >>>>>>>>>>> while others may not. >>>>>>>>>>>>>>> On Sep 23, 2026, at 6:33 AM, Chris Lohfink >>>>>>>>>>>>>>> <[email protected]> >>>>>>>>>>> wrote: >>>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by >>>>>>>>>>>>>>> contributors >>>>>>>>>>> with demonstrated expertise >>>>>>>>>>>>>>> - Must have produced similar patches in size, scope and area >>>>>>>>>>> unassisted and with minimal third-party guidance >>>>>>>>>>>>>>> I really don't like this one or its wording. Definitely too "the >>>>>>>>>>> peasants are getting uppity lets build a wall". Lets not let a >>>>>>>>>>> subjective >>>>>>>>>>> thing like demonstrated expertise (who decides that?) be if it's ok >>>>>>>>>>> or not. >>>>>>>>>>> Hold the same standards for code quality and process for it all. I >>>>>>>>>>> don't >>>>>>>>>>> want this to be: only people on the storage team in Apple can use >>>>>>>>>>> AI. >>>>>>>>>>>>>>> Chris >>>>>>>>>>>>>>> On Wed, Sep 23, 2026 at 12:16 AM <[email protected] <mailto: >>>>>>>>>>> [email protected]>> wrote: >>>>>>>>>>>>>>>> I agree with Stefan and think this is both a reasonable and >>>>>>>>>>> thoughtful proposal. >>>>>>>>>>>>>>>> Here are some things I like about it: >>>>>>>>>>>>>>>> – It outlines areas where LLM usage is unambiguously useful to >>>>>>>>>>>>>>>> the >>>>>>>>>>> project’s developers and users. >>>>>>>>>>>>>>>> – It defines a spectrum of recommendations and cautions. >>>>>>>>>>>>>>>> – The only prohibited areas are extremely narrow and say >>>>>>>>>>>>>>>> nothing >>>>>>>>>>> about code at all. >>>>>>>>>>>>>>>> Some in this thread are responding as if this proposal seeks to >>>>>>>>>>> prohibit or sharply limit use of LLMs. In fact, it’s one of the >>>>>>>>>>> most open >>>>>>>>>>> and welcoming I’ve seen for an OSS project of our size where many >>>>>>>>>>> are >>>>>>>>>>> adopting policies that simply ban them entirely. I’ve re-appended >>>>>>>>>>> the >>>>>>>>>>> proposal below my message as it seems to have been lost in threaded >>>>>>>>>>> replies, and would encourage folks to give it a second read. >>>>>>>>>>>>>>>> Some brief thoughts based on my own use of LLMs: >>>>>>>>>>>>>>>> – I find them fantastically useful for reviewing and >>>>>>>>>>>>>>>> identifying >>>>>>>>>>> problems that have slipped through review - primarily via Alex >>>>>>>>>>> Petrov’s >>>>>>>>>>> /deep-review skill, which I have running in a VM in a loop >>>>>>>>>>> executing over >>>>>>>>>>> every new commit in the project as of a few days ago. I will be >>>>>>>>>>> posting a >>>>>>>>>>> few hand-authored Jira tickets based on findings that appear >>>>>>>>>>> legitimate to >>>>>>>>>>> me. For now, the loop is posting them as issue drafts for my own >>>>>>>>>>> review on >>>>>>>>>>> my personal fork which you can find here: >>>>>>>>>>> https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug >>>>>>>>>>>>>>>> – They’re great for enabling use of model checkers and formal >>>>>>>>>>> methods where such work would have previously been prohibitively >>>>>>>>>>> expensive, >>>>>>>>>>> such as Blake’s work on a TLA+ proof of aspects of Mutation >>>>>>>>>>> Tracking and >>>>>>>>>>> Benedict/Fedor’s work on a machine-checkable proof of the Accord >>>>>>>>>>> protocol >>>>>>>>>>> in Lean. >>>>>>>>>>>>>>>> – They are stunning for allowing me to experiment with ideas >>>>>>>>>>>>>>>> that >>>>>>>>>>> would have otherwise been a summer internship’s scope of work. Some >>>>>>>>>>> examples include an io_uring prototype, exploring the impact of >>>>>>>>>>> page-aligned compressed chunk sizes, an API shim bridging the 3.x >>>>>>>>>>> and 4.x >>>>>>>>>>> Java Drivers, and potential enhancements to Zstandard. >>>>>>>>>>>>>>>> – And they shine when given grunt-work that is critical to the >>>>>>>>>>> project but a miserable labor for humans, such as triaging, >>>>>>>>>>> reproducing, >>>>>>>>>>> and root-causing flaky tests, which David Capwell now has running >>>>>>>>>>> in a loop >>>>>>>>>>> to help us improve CI stability in the project. >>>>>>>>>>>>>>>> I never thought I’d be so positive on what’s possible via >>>>>>>>>>>>>>>> language >>>>>>>>>>> models a year ago. At the same time, I also agree that they present >>>>>>>>>>> challenges and risks that can be managed through thoughtful >>>>>>>>>>> discussion and >>>>>>>>>>> policy. Some of the concerns that I think are important to guard >>>>>>>>>>> against >>>>>>>>>>> include: >>>>>>>>>>>>>>>> – Asymmetry of effort between author and reviewers: As >>>>>>>>>>> token-generating machines, LLMs can generate diffs of extraordinary >>>>>>>>>>> size >>>>>>>>>>> very rapidly. /deep-review is great for chewing through diffs and >>>>>>>>>>> identifying defects. But it should be used by the contributor >>>>>>>>>>> themselves to >>>>>>>>>>> identify issues – not to replace the role of the reviewer with more >>>>>>>>>>> electricity. The role of the reviewers extends beyond identifying >>>>>>>>>>> and >>>>>>>>>>> highlighting defects. It encompasses architecture, harmony with the >>>>>>>>>>> existing codebase, thinking ahead to future evolution of the >>>>>>>>>>> project, and >>>>>>>>>>> replicates context on the project as new code is committed. These >>>>>>>>>>> functions >>>>>>>>>>> cannot be automated away. >>>>>>>>>>>>>>>> – Hesitancy of authors to engage manually with code they have >>>>>>>>>>> generated: This is not specific to Cassandra, but it is a behavior >>>>>>>>>>> that I >>>>>>>>>>> have seen in several “highly-electric” projects. There’s a bimodal >>>>>>>>>>> tendency >>>>>>>>>>> toward code that is entirely generated or entirely human-authored - >>>>>>>>>>> but it >>>>>>>>>>> is rare for someone to prepare an AI-authored patch to take an >>>>>>>>>>> offramp and >>>>>>>>>>> spend a significant amount of time refining the work by hand in an >>>>>>>>>>> IDE. >>>>>>>>>>> This hesitancy toward human participation in authorship of >>>>>>>>>>> LLM-generated >>>>>>>>>>> code is very concerning to me. >>>>>>>>>>>>>>>> – Harmony with the existing codebase: Due to the tunnel-vision >>>>>>>>>>>>>>>> of >>>>>>>>>>> context windows, LLMs are generally unaware of conventions and norms >>>>>>>>>>> present in codebases and very frequently reinvent concepts in a >>>>>>>>>>> generation >>>>>>>>>>> turn to suit a goal without view of the project’s overall >>>>>>>>>>> architecture. >>>>>>>>>>> This results in a profusion of messy and duplicated concepts that >>>>>>>>>>> gradually >>>>>>>>>>> sprawl about a codebase. >>>>>>>>>>>>>>>> Again, none of these are grounds for prohibition of usage of >>>>>>>>>>> language models in developing the project. They’re just problems we >>>>>>>>>>> need to >>>>>>>>>>> bear in mind and guard against – and I think the proposal is >>>>>>>>>>> designed to do >>>>>>>>>>> just that. >>>>>>>>>>>>>>>> I’m thrilled by the potential of LLMs to improve Apache >>>>>>>>>>>>>>>> Cassandra >>>>>>>>>>> and we already see it happening through a vast number of issues >>>>>>>>>>> that are >>>>>>>>>>> being reported and fixed. But there’s also danger in taking ATVs >>>>>>>>>>> down a >>>>>>>>>>> hiking trail full of people. >>>>>>>>>>>>>>>> Regarding the prohibition on prose, I’ll simply say: I recently >>>>>>>>>>> found myself in a scenario where I found a Claude-authored document >>>>>>>>>>> so >>>>>>>>>>> inscrutable that I piped it back into a model, directed it to >>>>>>>>>>> rewrite it in >>>>>>>>>>> ASD-STE100, read it myself, and responded based on the >>>>>>>>>>> summarization. As a >>>>>>>>>>> humanities grad, this is probably the worst language crime I have >>>>>>>>>>> committed. But it was in response to language that was itself so >>>>>>>>>>> idiosyncratic that it was unreadable to me in its original form. I >>>>>>>>>>> hope >>>>>>>>>>> this never happens in the Apache Cassandra project. >>>>>>>>>>>>>>>> I’ll close with a quote from an excellent article written by >>>>>>>>>>>>>>>> Colin >>>>>>>>>>> Breck, an engineer who works on large-scale data systems: >>>>>>>>>>> https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/ >>>>>>>>>>>>>>>> Colin wrote: >>>>>>>>>>>>>>>>> I don’t want to live in a world where you use AI to summarize >>>>>>>>>>> something important into unreadable text, and then I use AI in an >>>>>>>>>>> attempt >>>>>>>>>>> to decipher it. I want to hear you, imperfections and all. I want >>>>>>>>>>> your >>>>>>>>>>> interpretation of aesthetics, beauty, quality, relationship, time. >>>>>>>>>>> I want >>>>>>>>>>> to know how you feel. I want you to cut through and tell me what >>>>>>>>>>> really >>>>>>>>>>> matters. >>>>>>>>>>>>>>>>> Intentional writing will likely become more valuable. People >>>>>>>>>>>>>>>>> who >>>>>>>>>>> write, and write to think, to think deeply and carefully, or to >>>>>>>>>>> create, to >>>>>>>>>>> share, or to capture something important without explicitly >>>>>>>>>>> expressing it >>>>>>>>>>> will continue to write and produce original work. The people who >>>>>>>>>>> never were >>>>>>>>>>> writers will use AI to produce lots of text. >>>>>>>>>>>>>>>> I hope that our culture can remain one of intentional writing >>>>>>>>>>>>>>>> and >>>>>>>>>>> intentional engineering. I enjoy reading the voice of the author in >>>>>>>>>>> comments, code, and tickets in Cassandra – the different ways we use >>>>>>>>>>> language based on where we grew up and how we learned English, the >>>>>>>>>>> translated idioms from our various backgrounds, and terse comments >>>>>>>>>>> that >>>>>>>>>>> recognize the difference between code whose function is obvious and >>>>>>>>>>> what >>>>>>>>>>> warrants genuine exposition. When I read code in Cassandra, it’s a >>>>>>>>>>> delight >>>>>>>>>>> to recognize the author based on their writing style before >>>>>>>>>>> flipping on >>>>>>>>>>> `git annotate` to reveal the origin. >>>>>>>>>>>>>>>> I’d encourage folks to re-read the original proposal below. It >>>>>>>>>>>>>>>> is >>>>>>>>>>> very permissive. The guidance strikes me not just as reasonable, but >>>>>>>>>>> genuinely important to maintaining the health of the project. >>>>>>>>>>>>>>>> – Scott >>>>>>>>>>>>>>>> ===== >>>>>>>>>>>>>>>> Encouraged: >>>>>>>>>>>>>>>> - Reviewing and otherwise validating human-authored patches >>>>>>>>>>>>>>>> before >>>>>>>>>>> submission >>>>>>>>>>>>>>>> - Debugging, diagnosing etc >>>>>>>>>>>>>>>> Permitted: >>>>>>>>>>>>>>>> - Generating or modifying tests, scripts, tooling or any other >>>>>>>>>>> non-user facing changes >>>>>>>>>>>>>>>> - Minor changes to human-authored patches that are carefully >>>>>>>>>>> reviewed by the author >>>>>>>>>>>>>>>> Restricted: >>>>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by >>>>>>>>>>>>>>>> contributors >>>>>>>>>>> with demonstrated expertise >>>>>>>>>>>>>>>> - Must have produced similar patches in size, scope and area >>>>>>>>>>> unassisted and with minimal third-party guidance >>>>>>>>>>>>>>>> - Core code changes made by LLM require an additional reviewer >>>>>>>>>>>>>>>> - LLM review is not a substitute for human review, and must be >>>>>>>>>>>>>>>> used >>>>>>>>>>> only to augment a complete and independent human understanding of >>>>>>>>>>> the patch. >>>>>>>>>>>>>>>> Prohibited: >>>>>>>>>>>>>>>> - All public prose must be human authored. This includes inline >>>>>>>>>>> comments, docs, posts to Jira etc. >>>>>>>>>>>>>>>> All LLM generated changes MUST be disclosed: >>>>>>>>>>>>>>>> - Outlined to any reviewer; >>>>>>>>>>>>>>>> - Summarised in the commit message; >>>>>>>>>>>>>>>> - Large blocks or files must be individually marked with some >>>>>>>>>>>>>>>> agreed >>>>>>>>>>> message like "created by <some AI>" >>>>>>>>>>>>>>>> ===== >>>>>>>>>>>>>>>>> On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected] >>>>>>>>>>> <mailto:[email protected]>> wrote: >>>>>>>>>>>>>>>>> On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected] >>>>>>>>>>> <mailto:[email protected]>> wrote: >>>>>>>>>>>>>>>>>> Restricted: >>>>>>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by >>>>>>>>>>> contributors with demonstrated expertise >>>>>>>>>>>>>>>>>> - Must have produced similar patches in size, scope and area >>>>>>>>>>> unassisted and with minimal third-party guidance >>>>>>>>>>>>>>>>> I am -1 on this. This sounds like gate keeping attempt. It >>>>>>>>>>>>>>>>> narrowly >>>>>>>>>>> limits the pool to a few people on the project that have >>>>>>>>>>> historically >>>>>>>>>>> contributed to certain parts of the codebase. This policy will >>>>>>>>>>> prohibit >>>>>>>>>>> skilled software engineers with domain expertise from proposing LLM >>>>>>>>>>> assisted changes simply because they have not contributed to the >>>>>>>>>>> project. >>>>>>>>>>> This is unrealistic and a net negative for the project to attract >>>>>>>>>>> talent >>>>>>>>>>> and grow our community. >>>>>>>>>>>>>>>>>> - Core code changes made by LLM require an additional >>>>>>>>>>>>>>>>>> reviewer >>>>>>>>>>>>>>>>> Can you be more precise what is this in addition to? How many >>>>>>>>>>>>>>>>> total >>>>>>>>>>> reviewers do you expect and what is the purpose of additional >>>>>>>>>>> reviewer? and >>>>>>>>>>> why? >>>>>>>>>>>>>>>>> Taking a step back - what are you trying to solve here? >>>>>>>>>>>>>>>>> Dinesh
