+1 On Sat, Sep 26, 2026 at 9:21 AM Jon Haddad <[email protected]> wrote:
> +1 to simple straightforward guidelines > > > > On Thu, Sep 24, 2026 at 1:01 PM Blake Eggleston <[email protected]> > wrote: > >> I propose the following guidelines: >> >> >> 1. Contributors must have a comprehensive understanding of the >> patches they contribute >> 2. Reviewers must have a comprehensive understanding of the patches >> they approve >> 3. Reviewers must have experience with the affected systems >> commensurate with the risk and complexity of the patch. >> >> >> This makes the expectations that have been mostly implicit explicit, >> doesn’t gate keep on the contributor side, doesn’t tell people how to do >> their work, and gives the subject matter expert, the reviewer, the latitude >> to decide what standards make sense for the patch they’re reviewing. >> >> >> On Thu, Sep 24, 2026, at 12:01 PM, Josh McKenzie wrote: >> >> Apache Cassandra was fundamentally undeployable for four years between >> Nov 2015 - 2019. >> >> CASSANDRA-8099 was a maximal manifestation of a specific approach to >> engineering and calendar constraints we've seen time and again on the >> project; I don't want us to conflate things here. That was a herculean >> monolithic body of work performed in inhuman conditions (in vim!) that was >> ultimately so invasive, all the unit tests in the code-base were commented >> out and Jake and I spent a grueling 1.5-2 months hand-rewriting basically >> all the unit tests in that code-base to get things to even build and run, >> much less pass. >> >> Massive blast radius changes that are un-sustainably complex, >> under-tested, where we don't property, fuzz, check coverage, check >> complexity, or A/B compare against a known good system (i.e. pre/post) >> correctness testing are going to destabilize the database at any time under >> any regime of tooling. We certainly could speed-run our way back into >> destabilization with LLM's as they are a force multiplier for both the good >> and the bad of one's engineering practices, but that's a solvable problem >> by better defining what our bars of quality are (Definition of Done >> anyone?) and holding ourselves accountable to delivering at that bar. >> >> On Thu, Sep 24, 2026, at 2:40 PM, David Capwell via dev wrote: >> >> >> if you really want to pursue it I would ask that we do it offline to >> avoid polluting an already busy conversation >> >> People are directly responding saying that they feel discrimination >> currently and that the policy tries to codify that discrimination, so I >> feel its 100% on topic. This thread has presented 0 evidence that LLM usage >> has lowered the quality of contributions merged and has so far been vibes >> and feeling; I have yet to see any evidence to justify such discrimination >> so I will keep pushing back until such evidence is presented so we can have >> a informed debate. >> >> namely how we handle shallow and localised bug fixes. I would be happy >> adding a clear entry to the “Permitted” section for this. No doubt there >> are many refinements needed to the Restricted text as well, that might also >> capture some of your concerns. >> >> I will not sign off on cherry picking areas where "safe" to use a tool; >> so no it does not capture my concerns. >> >> >> I don't know if everyone remembers, but ten years ago Cassandra was full >> of serious correctness and stability issues. Despite developing it, I would >> not have run it myself or recommend that anyone use it. We have dug >> ourselves out of that hole, but it took years of discipline and effort, and >> we're still (deservedly) recovering our reputation. >> >> Let's use this new technology to improve the quality of our >> contributions, not squander our hard-earned gains in the name of speed. It >> will be hard to recover our reputation a second time. >> >> I do recall the 3.x line and put in a significant amount of effort to >> harden it. There were behaviors I noticed after joining Cassandra that I >> feel directly contributed to 3.0 and the decade of catch up; behaviors that >> still linger in parts of the community today. >> >> As I look on trunk and look at committed code and trace back to PRs and >> JIRA I see the following: >> >> - large patches approved without comments >> - 0 evidence that tests were run >> >> I then look at our CI and see tests failing for months. As you start to >> triage you start to see some of them show real issues; yet they linger for >> months not being addressed... when CI is unstable it takes a lot of effort >> to triage "did my patch break the test", and I have seen time and time >> again people do not put in that effort, and shrug off as "its just a flakey >> test"; then our CI failure rate grows. >> >> Non of this has anything to do with LLMs but LLMs running in this >> environment is far more dangerous as there are not checks in place to "hold >> the bar". I am all for raising the bar universally; expecting both humans >> and LLMs to match that bar. >> >> I’d support something that boils down to roughly this: >> >> 1.) 2 committers must understand an LLM-assisted change before it commits. >> (Perhaps separately we can explore the question of why we haven’t added any >> new committers to the core project for about a year. I’m also still not >> entirely sure if it’s acceptable within our guidelines for a committer to +1 >> a patch after delegating review.) >> >> 2.) Patch authors must demonstrate enough understanding to discuss their own >> patch, whether or not parts of it are generated by an LLM. >> >> 3.) The “Assisted-by” tag should be used to indicate any non-trivial LLM >> usage in the generation of a patch, just like we have used Co-authored-by >> historically. >> >> 4.) Comments and other things that aren't the actual code (but could sow >> confusion) should be held to the same standard we'd expect from a human >> writer. If we don’t yet agree on that standard, we can formalize enough of >> it to guide both humans and LLMs. >> >> I can get behind this proposal but i would tweak it as 1/2 i don't think >> really need to special case LLM usage >> >> 1. 2 committers must understand the change before it commits. >> 2. Patch authors must demonstrate enough understanding to discuss >> their own patch >> >> Nothing about those 2 need to be scoped to LLM usage and honestly matches >> most PMCs I have talked to understanding of our bar (as Benedict pointed >> out, the actual wording could be interpreted to allow rubber stamping from >> committer) >> >> As for 3 I am cool with this. ASF recommends the same (it says >> Generated-by but that discount's the human's effort) as its useful for >> audits and tooling. Having Assisted-by tag should not imply anything >> about the committed patch as it should have gone through the same bar we >> all expect; its just for tools auditing. >> >> On Sep 24, 2026, at 10:37 AM, Aleksey Yeshchenko via dev < >> [email protected]> wrote: >> >> Meant "doing away with", sorry. Non-native speaker with a headache here. >> Thanks Caleb for spotting. >> >> On 24 Sep 2026, at 17:35, Aleksey Yeshchenko via dev < >> [email protected]> wrote: >> >> P.S. I assume it's obvious from the text above that I don't believe that >> getting away with human code review is a viable option. >> >> >> >> On 24 Sep 2026, at 17:37, Štefan Miklošovič <[email protected]> >> wrote: >> >> Good call on checkerframework, we even have a patch for it. Work of >> Jacek Lewandowski. We might just drive it to completion. Using AI for >> finishing it would be quite ironic. >> >> (1) https://github.com/apache/cassandra/pull/2370 >> >> On Thu, Sep 24, 2026 at 6:19 PM Jon Haddad <[email protected]> >> wrote: >> >> >> There are some really good points being brought up about stability of the >> codebase, maintainability, quality of reviews, correctness bugs, and I >> agree with all of them. I think it would be helpful to take a step back >> and consider how those bugs got there in the first place, how they were >> fixed, and what we could do to further advance the codebase so they don't >> creep back. LLMs can be used either with very tight guardrails, or in >> what's effectively YOLO mode, and there's a big difference in the quality >> of the results you get. >> >> One thing to keep in mind, a lot of the initial code in C* was added >> without comprehensive testing. I hope we can all agree that it's a lot >> easier to break code that doesn't have high quality tests. During the code >> freeze, a lot of people people spent several years relentlessly finding and >> fixing bugs. This was probably a pretty frustrating time for anyone who was >> focused on fixing other people's bugs when they wanted to build features. >> I think we should recognize the effort here and appreciate the foundation >> that the project stands on now. I can understand how anyone involved with >> this effort would be apprehensive about seeing years of their life swept >> away by an agent that was driven by goal seeking to remove all the tests >> that it broke instead of fixing them. >> >> When I picked up the work to improve cursor compaction, the first thing I >> asked myself was how can I make sure I don't break this? How do I even >> know it works properly? There were some tricky parts to the code, and I >> really didn't want to come in and immediately break stuff. That's why I >> started with an entire patch dedicated to adding test infra to it. 90% of >> the patch was tests, and in my other cursor patches, it remains *at least* >> 80% of my patches. It was a *lot* faster to add almost 10K lines of tests >> that handled a byte for byte differential testing paired with harry to find >> over 30 bugs that caused cursor to corrupt results. Range tombstones alone >> were at least a dozen bugs, but I also found issues with static columns, >> reverse ordering, etc. Randomizing schemas and data in burn tests to >> generate different shapes of data, to ensure they all result in the same >> output at the end. JMH tests to ensure there weren't performance >> regressions, hours of profiling. These were all a *lot* easier to do with >> the LLM helping me out. In the process I've found bugs that have been >> lingering in the codebase for years. >> >> That's a long story, but hopefully we all agree that having comprehensive >> tests is a great way to ensure that both humans and LLMs don't break things >> that are working. >> >> The lesson: we need to keep improving our testing. Everything that we >> touch, should be left in a better state than how we found it with regard to >> test coverage. >> >> Test coverage isn't everything though, there's always little subtle bugs >> that don't get found in testing, that can slip in despite our best >> efforts. It's debatable if humans will be as good as agents for coding in >> the long term, for spotting small defects. I sincerely doubt it. For the >> time being though, we still have people involved. It's probably a good time >> to start using more static analysis tools to identify problematic code and >> to add this to CI. Dmitry had a suggestion recently for checkerframework >> to detect leaking contexts, a problem he spotted when reviewing my branch. >> It would be great to have that integrated into our CI and dev workflow so >> we can simply avoid an entire class of bugs. >> >> There's also PMD, which is excellent for finding code that can be hard to >> understand. I *highly* suggest you all run PMD to analyze for cognitive >> complexity and high npath scores. This was made popular by the folks at >> Sonar and I've found it to be an excellent feedback mechanism for >> structuring code. The default max they set is 15, which is the point where >> it starts to become difficult to verify something works without making a >> massive investment. We've got areas in the codebase that are in the >> hundreds, and some parts even higher. These have been contributed by >> humans, and are all high risk points for both humans and agents to start >> messing around with. They're also in some fairly critical areas that are >> very likely to break, so I understand why people would not want an agent >> anywhere near it. >> >> Unfortunately, it's not an easy problem to address. There's so many >> places where the code is structured in a way that has so many branches, so >> many conditions, that it's effectively impossible for a human to >> understand, creating a fear of messing around in it. There's plenty of >> areas that deserve extreme scrutiny, and we should be careful of what we >> add, whether it's human or agent. >> >> The codebase today requires a high degree of internal knowledge to >> navigate. There's land mines everywhere. We should be looking to make >> conscious improvements by moving the code forward, so it's easier to make >> changes to small, well tested components with minimal side effects. Not >> making it harder for people to use the tools that aid in that process. >> >> Here's what we could do to achieve the underlying goal of not breaking >> the DB: >> >> Add cognitive complexlity and npath via PMD as a feedback mechanism. >> >> Code that's hard to understand is hard to review. It's also hard to >> test. Let's break down the complex code so more people can contribute, >> safely. >> >> Add checkerframework to our tooling, >> >> Properly annotate the codebase for it and reduce the surface area that >> things can break. Less brittle codebase = we can move faster. >> >> Use jacoco to find areas of the codebase with poor testing. >> >> Let's improve the test coverage there, LLMs are great for this. We have >> a ton of static tests, these can become more dynamic, parameterized, and >> leverage harry. >> >> Refactor parts of the codebase that have high cognitive complexlity and >> NPath scores. >> >> This should be lowered over time to meet some high watermark, say 25 >> maximum, although I'd prefer 15 which is where the Sonar folks settled. >> >> Move forward moving the codebase to a more modular structure >> >> We've talked about Gradle on and off - but it can really be a huge help >> with incremental, modular builds. This is pretty easy to do with an agent >> and we could have it done in a couple days. >> >> Enforce boundaries with ArchUnit >> >> If we want to enforce certain code boundaries, this is the way to do it. >> Should not be part of manual review. >> >> Add LLM review for all incoming PRs before a human >> >> The goal here is to automate the initial part of the review process that >> reviewers should spot, and raise the bar for the initial contribution. >> When the code gets reviewed by a human, it should already have passed a >> large variety of initial checks. This should shorten the review cycle and >> result in higher quality patches. I've had Claude reviewing all my PRs in >> my personal projects for a while now and it consistently gives great >> feedback that I almost always incorporate. >> >> In my ideal world, we'd also auto-format all code >> >> Consistent formatting throughout the codebase would be amazing, but >> that's just one man's dream. >> >> Hopefully there's at least a couple things in this list we could move >> forward with in the short term, as it'll help improve the code quality >> regardless of how it's created. >> >> Jon >> >> https://checkerframework.org/manual/#aliasing-leaking-contexts >> https://www.sonarsource.com/docs/CognitiveComplexity.pdf >> https://pmd.github.io/pmd/pmd_rules_java_design.html >> >> >> >> >> >> On Thu, Sep 24, 2026 at 7:38 AM C. Scott Andreas <[email protected]> >> wrote: >> >> >> From Benedict: >> >> “I don't know if everyone remembers, but ten years ago Cassandra was full >> of serious correctness and stability issues. Despite developing it, I would >> not have run it myself or recommend that anyone use it. We have dug >> ourselves out of that hole, but it took years of discipline and effort, and >> we're still (deservedly) recovering our reputation.” >> >> Expanding on this point for those who may not have been active in the >> project at this time — >> >> Apache Cassandra was fundamentally undeployable for four years between >> Nov 2015 - 2019. The database literally lost data if you ran a read-only >> SELECT query ordered descending (C-14513, C-14515). If you haven’t read >> these tickets before, please take a moment to do so. >> >> It took years of careful work via property-based testing, fuzzing, and >> deterministic simulation to restore Cassandra’s status as a usable system >> of record. Once 14513 and 14515 were identified, nearly 30 additional >> critical data loss and incorrect response bugs were identified. >> >> It is essential for the project’s future that we don’t regress to this >> state chasing AI-generated features motivated by fear. The fact that >> examples cited in this thread which boast shiny features but have critical >> shortcomings unknown to their author supports this argument. >> >> The most common path for large corpuses of AI-generated software is >> elation and reveling in a feature matrix, followed by abandonment. >> >> I endorse this point: >> >> “Let's use this new technology to improve the quality of our >> contributions, not squander our hard-earned gains in the name of speed. It >> will be hard to recover our reputation a second time.” >> >> Patrick, I don’t want your note regarding a TCM issue to go unaddressed. >> Please file a Jira ticket and the patch if you like. I can’t comment on the >> patch as I haven’t seen it, but together we will solve the problem. >> >> – Scott >> >> On Sep 24, 2026, at 4:01 AM, Benedict Elliott Smith <[email protected]> >> wrote: >> >> Hi Patrick, >> >> As I mentioned in my reply to David, I would be happy to create a carve >> out for shallow and localised bug fixes in the "Permitted" section. Would >> this alleviate some of your concerns regarding your ability to contribute >> to the project? >> >> I appreciate your pointing out Ferrosa's Accord implementation however, >> as it is a *great* example of the problems we're leaping into. I took a >> look, and within about 30s found that the protocol is fundamentally >> incorrect, having failed to address CASSANDRA-18365. This is despite >> claiming to be tested with Jepsen that should in principle find this fault. >> I followed up by using Claude to interrogate the implementation further, >> and immediately found other serious correctness issues. >> >> I use LLMs daily now to help facilitate Accord development, and while >> they are powerful they are NOT able to author the code themselves, even >> when building upon a strong human-authored foundation. >> >> I don't know if everyone remembers, but ten years ago Cassandra was full >> of serious correctness and stability issues. Despite developing it, I would >> not have run it myself or recommend that anyone use it. We have dug >> ourselves out of that hole, but it took years of discipline and effort, and >> we're still (deservedly) recovering our reputation. >> >> Let's use this new technology to improve the quality of our >> contributions, not squander our hard-earned gains in the name of speed. It >> will be hard to recover our reputation a second time. >> >> >> On 2026/09/23 19:16:17 Patrick McFadin wrote: >> I was waiting for this moment to hit our project and I'm glad we're here. >> I >> am deeply concerned for our project and its future, as we have >> increasingly >> made it difficult to contribute. I had hoped that this new era of >> software tools powered by AI would expand the project's reach and bring >> more diverse thoughts and ideas. This policy proposal is the exact >> opposite >> of what we need. We have been sitting on a Cassandra 6 release alpha for >> months. We need to accelerate and embrace new ways of being or be left >> behind. As I read that policy, my first and gut level reactions: >> - It comes across as elitist and class protectionism. Committer should not >> be special but this proposal makes that designation even more sacred. >> - It signals that our project is so fragile that only a few people "Really >> understand it" That's some SQLite vibes right there. >> - Trying to fix a problem that doesn't exist >> Sadly, i think this policy change would also exclude a lot of comitters. >> We aren't alone in this moment. The Linux project just went through >> this. You can find the thread with a simple Google, but similar hard >> feelings were being expressed "AI is going to ruin our project!", "The >> unwashed masses are going to contribute terrible code!", "We have to >> protect our precious status as Linux maintainers!" Linus being Linus, was >> deeply invloved and they adopted a super simple statement that covers all >> bases. Human or Human using AI. “You are expected to understand and to be >> able to defend everything you submit.” Love that. >> In the larger picture, I'll restate. I'm worried for our project. In late >> 2025(Opus 4.5 IYKYK), early 2026, AI coding LLMs turned a real corner and >> in the hands of somebody that knows how to build software, this tool is >> like jet fuel. Here's some examples of new projects being hyper fueled by >> AI coding tools. >> Apache Iggy - Complete rust replacement of kafka. Crazy fast velocity >> Turso - Rust re-write of SQLite >> Bun - Rust re-write of itself from Zig. >> Think this couldn't happen to us? Already has: >> https://github.com/ferrosadb/ferrosa. Ben is using it to power his own >> startup, but it was him alone using a ton of local AI coding agents. He >> even implemented Accord. Yeah... >> The cracks are already starting to show. There is a black market economy >> of >> Cassandra patches happening now. Not going to name names or call people >> out, but there are fixes and optimizations living in branches outside of >> the Cassandra project. Why? I'll use myself as an example. I fixed a nasty >> bug I ran into with TCM a few weeks ago. Wrote the tests. It passes CI and >> lives in my personal branch. I'm sitting here really wondering if I want >> to >> go through the ritual humiliation of being roasted for using AI to fix it. >> Me. I am worried about contrinuting code the Cassandra. What the hell does >> that say? >> I have my CQLite project that I've been doing a release around once a >> month. I would love to donate that to the Cassandra project but I wouldn't >> if it essentially killed any progress. >> My larger counter proposal would be to: >> - Adopt the “You are expected to understand and to be able to defend >> everything you submit.” approach the Linux project has adopted. >> - Loosen up the contributor process and our worry on trunk. Let 1000 >> flowers bloom and bring it in. >> - And finally, to give some people more peace of mind and open more doors, >> adopt what other projects have done and provide more pluggability. Let new >> ideas have an easy place to connect. >> We are at a fork in the road. What are we going to do? And then I have to >> ask myself, what am I going to do as a contributor? >> Patrick >> On Wed, Sep 23, 2026 at 6:16 AM Blake Eggleston <[email protected]> >> wrote: >> >> I’m not necessarily opposed to having a policy, but so far we have some >> specific proposals addressing a problem statement that’s very nebulous. >> What is the community failing to do on its own that we’re trying to >> correct >> with policy? What outcomes are we trying to create or prevent? Having some >> examples and specific problems to discuss would help focus the >> conversation. >> >> On Wed, Sep 23, 2026, at 4:34 AM, Shailaja Koppu via dev wrote: >> >> Benedict, >> Thanks for clarifying. My concern still remains. This criteria would be >> difficult to define and apply consistently. What counts as “similar” scope >> or area, “mostly correct,” or sufficiently independent work? More >> importantly, how do we prevent such vague criteria from creating an >> informal hierarchy where some contributors work is routinely accepted >> while >> others is routinely rejected? >> If the intent is to limit AI-assisted code changes to Cassandra >> contributors, or to contributors who have previously worked in that >> component without AI, that would at least be clear and enforceable. >> >> On Sep 23, 2026, at 12:01 PM, Benedict Elliott Smith < >> >> [email protected]> wrote: >> >> Core code changes >> Chris: Do you object to the first or second line you quote? Because the >> >> first line is effectively motivation for the second line, and can be >> removed (or more clearly combined). If it’s the second line, then I do not >> think this is an unreasonable expectation, and we can get into a proper >> debate about it. >> >> Shailaja, since you only snipped the first sentence, your concerns might >> >> also be mostly answered by this clarification? “Minimal third-party >> guidance” implies you have some concerns about the second line, but all of >> our policies have some ambiguity because legalese is even worse. I don’t >> think the ambiguity here would be challenging to navigate though we can >> certainly refine it. This specific snippet is meant to convey an >> expectation that a contributor has autonomously produced patches of >> similar >> scope that were mostly correct, so that they have demonstrated the level >> of >> understanding necessary to guide another party to a successful patch (i.e. >> an LLM in this case). >> >> On 2026/09/23 10:54:16 Benedict Elliott Smith wrote: >> >> Thanks everyone for your input so far. I’ll respond in brief to the >> >> main themes, in (mostly) separate emails so they can each have their own >> debate chain. >> >> Should we have a policy (Blake/Josh*/Jon/Dinesh) >> I think we would all agree that LLMs represent the biggest change to >> >> this community (and software more generally) since its inception, and we >> all now have enough experience with the technology to have formed opinions >> about how it is best managed. We also evidently have not all arrived at >> the >> same conclusions. In this situation, it would be an abdication of our >> responsibilities as a management committee to not agree *some* policy. >> >> I intend to conduct straw polls as the discussion evolves, so if you >> >> prefer an alternative policy - or modifications to this policy - I would >> encourage you to make those alternative proposals. >> >> *Veto/Consensus (Josh) >> It was fair to call out my poor use of language on this topic, so let >> >> me rephrase a little. The community is built on consensus, and work should >> not be merged when there are outstanding concerns to address. The explicit >> -1 should only be used rarely, because the prior expectation should >> prevent >> it ever being needed. I (and others) have outstanding concerns on LLM >> generated work that can only be addressed through this process right here, >> so to merge such work while maintaining the community’s consensus we must >> agree some policy. >> >> On 2026/09/23 09:58:27 Shailaja Koppu via dev wrote: >> >> I am strongly -1 on this >> - Core code changes made by LLM may only be proposed by contributors >> >> with demonstrated expertise >> >> That creates a new, subjective privileged class of contributors and >> >> turns a tool choice into an eligibility test. Who decides whether >> expertise >> has been “demonstrated,” what counts as “minimal third-party guidance,” >> and >> how could those judgments be applied consistently or fairly? >> >> Apache already has a better model, anyone may contribute, trust and >> >> additional repository privileges are earned transparently over time. The >> ASF describes its communities as flat, and says that newcomer ideas have >> as >> much input as those from original creators. We should not add a separate, >> informal hierarchy in which certain people may use common development >> tools >> while others may not. >> >> On Sep 23, 2026, at 6:33 AM, Chris Lohfink <[email protected]> >> >> wrote: >> >> - Core code changes made by LLM may only be proposed by contributors >> >> with demonstrated expertise >> >> - Must have produced similar patches in size, scope and area >> >> unassisted and with minimal third-party guidance >> >> I really don't like this one or its wording. Definitely too "the >> >> peasants are getting uppity lets build a wall". Lets not let a subjective >> thing like demonstrated expertise (who decides that?) be if it's ok or >> not. >> Hold the same standards for code quality and process for it all. I don't >> want this to be: only people on the storage team in Apple can use AI. >> >> Chris >> On Wed, Sep 23, 2026 at 12:16 AM <[email protected] <mailto: >> >> [email protected]>> wrote: >> >> I agree with Stefan and think this is both a reasonable and >> >> thoughtful proposal. >> >> Here are some things I like about it: >> – It outlines areas where LLM usage is unambiguously useful to the >> >> project’s developers and users. >> >> – It defines a spectrum of recommendations and cautions. >> – The only prohibited areas are extremely narrow and say nothing >> >> about code at all. >> >> Some in this thread are responding as if this proposal seeks to >> >> prohibit or sharply limit use of LLMs. In fact, it’s one of the most open >> and welcoming I’ve seen for an OSS project of our size where many are >> adopting policies that simply ban them entirely. I’ve re-appended the >> proposal below my message as it seems to have been lost in threaded >> replies, and would encourage folks to give it a second read. >> >> Some brief thoughts based on my own use of LLMs: >> – I find them fantastically useful for reviewing and identifying >> >> problems that have slipped through review - primarily via Alex Petrov’s >> /deep-review skill, which I have running in a VM in a loop executing over >> every new commit in the project as of a few days ago. I will be posting a >> few hand-authored Jira tickets based on findings that appear legitimate to >> me. For now, the loop is posting them as issue drafts for my own review on >> my personal fork which you can find here: >> >> https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug >> >> – They’re great for enabling use of model checkers and formal >> >> methods where such work would have previously been prohibitively >> expensive, >> such as Blake’s work on a TLA+ proof of aspects of Mutation Tracking and >> Benedict/Fedor’s work on a machine-checkable proof of the Accord protocol >> in Lean. >> >> – They are stunning for allowing me to experiment with ideas that >> >> would have otherwise been a summer internship’s scope of work. Some >> examples include an io_uring prototype, exploring the impact of >> page-aligned compressed chunk sizes, an API shim bridging the 3.x and 4.x >> Java Drivers, and potential enhancements to Zstandard. >> >> – And they shine when given grunt-work that is critical to the >> >> project but a miserable labor for humans, such as triaging, reproducing, >> and root-causing flaky tests, which David Capwell now has running in a >> loop >> to help us improve CI stability in the project. >> >> I never thought I’d be so positive on what’s possible via language >> >> models a year ago. At the same time, I also agree that they present >> challenges and risks that can be managed through thoughtful discussion and >> policy. Some of the concerns that I think are important to guard against >> include: >> >> – Asymmetry of effort between author and reviewers: As >> >> token-generating machines, LLMs can generate diffs of extraordinary size >> very rapidly. /deep-review is great for chewing through diffs and >> identifying defects. But it should be used by the contributor themselves >> to >> identify issues – not to replace the role of the reviewer with more >> electricity. The role of the reviewers extends beyond identifying and >> highlighting defects. It encompasses architecture, harmony with the >> existing codebase, thinking ahead to future evolution of the project, and >> replicates context on the project as new code is committed. These >> functions >> cannot be automated away. >> >> – Hesitancy of authors to engage manually with code they have >> >> generated: This is not specific to Cassandra, but it is a behavior that I >> have seen in several “highly-electric” projects. There’s a bimodal >> tendency >> toward code that is entirely generated or entirely human-authored - but it >> is rare for someone to prepare an AI-authored patch to take an offramp and >> spend a significant amount of time refining the work by hand in an IDE. >> This hesitancy toward human participation in authorship of LLM-generated >> code is very concerning to me. >> >> – Harmony with the existing codebase: Due to the tunnel-vision of >> >> context windows, LLMs are generally unaware of conventions and norms >> present in codebases and very frequently reinvent concepts in a generation >> turn to suit a goal without view of the project’s overall architecture. >> This results in a profusion of messy and duplicated concepts that >> gradually >> sprawl about a codebase. >> >> Again, none of these are grounds for prohibition of usage of >> >> language models in developing the project. They’re just problems we need >> to >> bear in mind and guard against – and I think the proposal is designed to >> do >> just that. >> >> I’m thrilled by the potential of LLMs to improve Apache Cassandra >> >> and we already see it happening through a vast number of issues that are >> being reported and fixed. But there’s also danger in taking ATVs down a >> hiking trail full of people. >> >> Regarding the prohibition on prose, I’ll simply say: I recently >> >> found myself in a scenario where I found a Claude-authored document so >> inscrutable that I piped it back into a model, directed it to rewrite it >> in >> ASD-STE100, read it myself, and responded based on the summarization. As a >> humanities grad, this is probably the worst language crime I have >> committed. But it was in response to language that was itself so >> idiosyncratic that it was unreadable to me in its original form. I hope >> this never happens in the Apache Cassandra project. >> >> I’ll close with a quote from an excellent article written by Colin >> >> Breck, an engineer who works on large-scale data systems: >> https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/ >> >> Colin wrote: >> >> I don’t want to live in a world where you use AI to summarize >> >> something important into unreadable text, and then I use AI in an attempt >> to decipher it. I want to hear you, imperfections and all. I want your >> interpretation of aesthetics, beauty, quality, relationship, time. I want >> to know how you feel. I want you to cut through and tell me what really >> matters. >> >> Intentional writing will likely become more valuable. People who >> >> write, and write to think, to think deeply and carefully, or to create, to >> share, or to capture something important without explicitly expressing it >> will continue to write and produce original work. The people who never >> were >> writers will use AI to produce lots of text. >> >> I hope that our culture can remain one of intentional writing and >> >> intentional engineering. I enjoy reading the voice of the author in >> comments, code, and tickets in Cassandra – the different ways we use >> language based on where we grew up and how we learned English, the >> translated idioms from our various backgrounds, and terse comments that >> recognize the difference between code whose function is obvious and what >> warrants genuine exposition. When I read code in Cassandra, it’s a delight >> to recognize the author based on their writing style before flipping on >> `git annotate` to reveal the origin. >> >> I’d encourage folks to re-read the original proposal below. It is >> >> very permissive. The guidance strikes me not just as reasonable, but >> genuinely important to maintaining the health of the project. >> >> – Scott >> ===== >> Encouraged: >> - Reviewing and otherwise validating human-authored patches before >> >> submission >> >> - Debugging, diagnosing etc >> Permitted: >> - Generating or modifying tests, scripts, tooling or any other >> >> non-user facing changes >> >> - Minor changes to human-authored patches that are carefully >> >> reviewed by the author >> >> Restricted: >> - Core code changes made by LLM may only be proposed by contributors >> >> with demonstrated expertise >> >> - Must have produced similar patches in size, scope and area >> >> unassisted and with minimal third-party guidance >> >> - Core code changes made by LLM require an additional reviewer >> - LLM review is not a substitute for human review, and must be used >> >> only to augment a complete and independent human understanding of the >> patch. >> >> Prohibited: >> - All public prose must be human authored. This includes inline >> >> comments, docs, posts to Jira etc. >> >> All LLM generated changes MUST be disclosed: >> - Outlined to any reviewer; >> - Summarised in the commit message; >> - Large blocks or files must be individually marked with some agreed >> >> message like "created by <some AI>" >> >> ===== >> >> On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected] >> >> <mailto:[email protected]>> wrote: >> >> On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected] >> >> <mailto:[email protected]>> wrote: >> >> Restricted: >> - Core code changes made by LLM may only be proposed by >> >> contributors with demonstrated expertise >> >> - Must have produced similar patches in size, scope and area >> >> unassisted and with minimal third-party guidance >> >> I am -1 on this. This sounds like gate keeping attempt. It narrowly >> >> limits the pool to a few people on the project that have historically >> contributed to certain parts of the codebase. This policy will prohibit >> skilled software engineers with domain expertise from proposing LLM >> assisted changes simply because they have not contributed to the project. >> This is unrealistic and a net negative for the project to attract talent >> and grow our community. >> >> - Core code changes made by LLM require an additional reviewer >> >> Can you be more precise what is this in addition to? How many total >> >> reviewers do you expect and what is the purpose of additional reviewer? >> and >> why? >> >> Taking a step back - what are you trying to solve here? >> Dinesh >> >>
