I think the main reason is that it's far less effort to create a comprehensive system for testing with an LLM than it is by hand. For example, with cursor compaction, we have parameterized, differential, fuzzed tests. I was able to go through several iterations and different ideas quite cheaply, in order to arrive where it is now. Being able to experiment with multiple paths forward is a huge general advantage, and on the side of testing it makes it a no brainer.
LLMs also don't complain when they need to make big revisions, or plumb things through that would otherwise be an annoyance. They don't mind the grunt work, and are damn good at it. Add a reviewer that runs PMD and Jacoco and tells you where the code is either too complex or untested AND can quickly refactor it, or put together a few hundred lines of test code is remarkable. Jon On Tue, Sep 29, 2026 at 1:56 PM Dinesh Joshi <[email protected]> wrote: > Jordan, thanks for the guiding principles. I think this is a short-enough > prose that I expect contributors to realistic read, digest and apply. > > One clarification though - why is the bar for LLM generated code higher > than human generated code? Why isn't the bar the same for both? Why does it > have to be a function of the thing that generated the code? > > Realistically, most developers are using LLM assistance is writing tests. > Like Caleb said, the cost of writing tests has fallen to zero. If anything, > I would say that the bar for tests should be higher regardless of who > generated the code. > > An unintended side effect of having a higher test bar for LLM generated > code is that it may incentivize contributors to hide / underplay the use of > LLMs. > > > > On Tue, Sep 29, 2026 at 1:04 PM Jordan West <[email protected]> wrote: > >> On Tue, Sep 29, 2026 at 12:54 Caleb Rackliffe <[email protected]> >> wrote: >> >>> @Jordan I agree with essentially everything you’ve said here. The only >>> exception is that I wouldn’t want to lower the testing bar (assuming we >>> agree on what that means) for non-LLM-authored patches. The cost of writing >>> tests has esssntially fallen to zero. >>> >> >> >> We’re absolutely aligned on that. If I implied otherwise that was an >> error on my part. My only intent is the bar should be even higher >> for LLM generated code than whatever the bar is for human generated code >> not that we lower the human generated bar we have or agree to in the >> future. Maybe we have to wordsmith that some to be more clear? >> >> >>> >>> On Sep 29, 2026, at 2:30 PM, Jordan West <[email protected]> wrote: >>> >>> >>> While I too want to lower the bar for contributors to join us, I am >>> happy to see something like the proposed policy and I don’t think reading a >>> short document is too much to ask when contributing to our large and >>> critical code base. We’ve had and have much larger barriers to entry than >>> that. One reason I’m excited about AI use in the project is I think it can >>> help us lower more of those. >>> >>> While I agree with the spirit of the proposed policy and some of what’s >>> in it I think as written it will have us back here often to re-discuss this >>> topic for a couple reasons. First, “strongly discouraged” will have >>> different meanings to different community members and as the policy >>> acknowledges it’s unenforceable so this will lead to debates based on the >>> ambiguities in the text. Second, all of us are on various points on a >>> spectrum on where we see AIs abilities today and where we see them going. >>> The policy is written to be a sort of blend of our current opinions on >>> where it is today so it seems likely we are back here today as our opinions >>> shift (I know my beliefs change daily to weekly these days in both >>> directions), capabilities change, and new consensus is needed. >>> >>> I propose we do have a document but one more I like of the “motivating >>> and guiding principles section” and that we rely on our other existing >>> policies and assumption of positive intent of contributors. I have proposed >>> some below. I am sure several here will find these too lenient given what I >>> presume to be where they fall on the AI adoption spectrum compared to me >>> and I respect that. Some days I am likely right there with you, others I’m >>> more bullish. You’ll find my proposal below does not encourage trying to >>> delineate parts of the codebase that can and cannot be contributed to with >>> an LLM but holds standards regardless. I encourage us to find ways to set >>> the bar for quality when LLMs are used now and in the future vs trying to >>> limit them based on today’s opinions and capabilities that are rapidly >>> changing. >>> >>> Some proposed ideas for some guiding principles with that in mind. It’s >>> likely not a complete list. >>> >>> * LLMs do not change accountability. You are ultimately accountable for >>> code produced or reviewed in your name. Whether hand written or by an LLM. >>> If you own an agentic process performing coding or review tasks you are >>> still responsible and accountable for what it produces or what actions it >>> takes. We as a community are accountable for understanding and being >>> knowledgeable about the software we provide to others. LLMs do not change >>> this. >>> >>> * Humans must be involved in the merging of code either by producing or >>> reviewing code and is subject to the accountability requirement above. >>> Nothing about using LLMs changes existing policies regarding committers >>> required to merge code, vetoes, or other voting procedures. >>> >>> * All LLM use must be attributed to both you and the harness, provider, >>> and model being used. The project makes no specific recommendations or >>> requirements regarding the toolchain used as long as the user has legal >>> access and provides attribution. >>> >>> * LLM generated code has a higher standard of automated testing than >>> human code, for which we have already adopted an incredibly high standard. >>> Proposed fixes or performance improvements must include runnable >>> demonstrations. Use of LLMs does not absolve the accountable contributor of >>> existing requirements such as providing a test plan in JIRA. LLM generated >>> code must be automatically linted to meet the projects code standards. LLM >>> generated code is not an excuse to ignore the projects existing policies on >>> code style. >>> >>> * It is strongly preferred that documentation and comments are not LLM >>> generated. We have found these to be of low quality. However, if done, the >>> aforementioned accountability remains with the contributor. “The LLM wrote >>> it” is not an acceptable dismissal of a review comment in any context. It >>> is recommended that documentation and comments continue to be human written >>> and optional LLM reviewed or edited with human supervision. >>> >>> On Tue, Sep 29, 2026 at 08:57 Caleb Rackliffe <[email protected]> >>> wrote: >>> >>>> @Josh I tried to define “directly generated” at the bottom, although >>>> there isn’t a proper footnote/link. I don’t think it matters at this point. >>>> There doesn’t appear to be any appetite for something in the middle, i.e. >>>> what I was attempting to do there. >>>> >>>> On Sep 29, 2026, at 10:00 AM, Josh McKenzie <[email protected]> >>>> wrote: >>>> >>>> >>>> Some questions that are still unclear to me after reading through this >>>> thread and the PR - and I assume a new contributor would be confused as >>>> well: >>>> >>>> Re: what qualifies as "Directly Generated" by an LLM: >>>> >>>> - If someone generates a full implementation and testing for >>>> something via an LLM then goes through line by line and cleans things up >>>> and makes changes, does that qualify as Directly Generated or not? >>>> - If they have fine-tuned a local model to comments in their own >>>> verbal style, is that strongly discouraged because an LLM generated it? >>>> What if they review it line-by-line? What if they write things by hand >>>> then >>>> have an LLM rephrase things and leave the LLM's final directly generated >>>> text in place? >>>> - What happens if 15% of the comments generated by the LLM are >>>> concise, clear, and only explain non-obvious "why's" of the code? >>>> Should a >>>> contributor go through and rephrase those lines in order to keep them >>>> from >>>> being directly generated? >>>> >>>> If I was a contributor looking for a project to start getting involved >>>> with baroque and bespoke rules would be incredibly off-putting to me. >>>> Honestly, the set of rules we have and social norms today are incredibly >>>> off-putting to many long-term contributors already who have a deep vested >>>> social and professional interest in the project succeeding. Who still would >>>> love to work technically on the project but are driven away by this >>>> culture. >>>> >>>> We're trying to hit a middle ground of not being too prescriptive but >>>> not leaving everything open to the interpretation of the reader which is >>>> just breeding more confusion. All in a space where the progress of the >>>> underlying tools is faster than anything I can recall in our field. >>>> Whatever policy we come up with now will probably be slightly outdated even >>>> by the time we ratify it unless it's incredibly high level and instead >>>> tries to codify our *values* and trust people to live up to them. >>>> >>>> Which I'd argue is exactly what Blake's simple proposal does. It's >>>> durable in the face of change and focuses on what's important to us and the >>>> community instead of engaging in pedantry and policing that just ends up >>>> confusing everyone further. >>>> >>>> On Tue, Sep 29, 2026, at 4:42 AM, Aleksey Yeshchenko via dev wrote: >>>> >>>> If anyone wants to follow along and/or add comments, we've created >>>> https://github.com/apache/cassandra/pull/5220 >>>> >>>> This is now quite qualitatively different from "Rust policy but without >>>> the committer exception for critical sections", I'm afraid. >>>> >>>> Watered down beyond what we discussed here and offline, and not what I >>>> and most folks who endorsed a Rust-like policy voted for. >>>> >>>> I'll make some edits to restore it to the shape we discussed last night. >>>> >>>> -- >>>> AY >>>> >>>> On 29 Sep 2026, at 00:17, Caleb Rackliffe <[email protected]> >>>> wrote: >>>> >>>> If anyone wants to follow along and/or add comments, we've created >>>> https://github.com/apache/cassandra/pull/5220 >>>> >>>> On Mon, Sep 28, 2026 at 2:45 PM Aleksey Yeshchenko via dev < >>>> [email protected]> wrote: >>>> >>>> It is a little confusing to keep track of the proposed diffs to Rust's >>>> policy. I think David is preparing a version with all the changes applied >>>> to it, so there is no ambiguity. >>>> >>>> How would we handle the “Non-critical” part of the experimental >>>> section? The policy exempts rust-lang members from that… does this mean >>>> we’d exempt committers but not non-committers. What’s the Cassandra analog >>>> of the non-critical section? >>>> >>>> >>>> The modified proposal removes that paragraph (about exemptions) >>>> altogether, thus treating all C* developers equally and allowing code >>>> generation for non-critical parts only for everyone. >>>> >>>> If you are still confused (which would be understandable - it is >>>> confusing), perhaps wait for David's doc, to make sure we are all on the >>>> same page wrt what's being proposed first. >>>> >>>> -- >>>> AY >>>> >>>> On 28 Sep 2026, at 20:16, Blake Eggleston <[email protected]> wrote: >>>> >>>> This is something I could support as well. >>>> >>>> 2 things: >>>> >>>> Our docs tend to be neglected. While ideally our docs would be 100% >>>> human generated, they’re mostly just not generated at the moment. While not >>>> ideal, I think relaxing the rust LLM policy as it relates to docs would be >>>> a net positive for users, provided they’re human reviewed and edited. >>>> >>>> How would we handle the “Non-critical” part of the experimental >>>> section? The policy exempts rust-lang members from that… does this mean >>>> we’d exempt committers but not non-committers. What’s the Cassandra analog >>>> of the non-critical section? >>>> >>>> On Mon, Sep 28, 2026, at 12:10 PM, Francisco Guerrero wrote: >>>> >>>> I've gone over the the Rust policy. I am in support of the Rust >>>> version with the tweaks proposed by Caleb. >>>> >>>> Best, >>>> - Francisco >>>> >>>> On 2026/09/28 18:45:03 Aleksey Yeshchenko via dev wrote: >>>> > Some last minute amends to the suggested policy's TL;DR, with Caleb's >>>> approval: >>>> > >>>> > - It’s fine to use LLMs to answer questions, analyze, distill, >>>> refine, check, suggest, review. >>>> > - LLMs work best when used as a tool to write better, not faster. >>>> > >>>> > "But not to create." bit is covered in detail by the full policy and >>>> is impossible to summarise well in two words. >>>> > >>>> > With other changes as outlined by Caleb in the quoted email, I would >>>> be happy to support this fine-tuned version of Rust's policy. >>>> > >>>> > -- >>>> > AY >>>> > >>>> > > On 28 Sep 2026, at 19:27, Caleb Rackliffe <[email protected]> >>>> wrote: >>>> > > >>>> > > To clarify, I would remove the "Experimental" tag and make that >>>> section apply to all contributors. (In other words, encourage attribution, >>>> quality, and human decision-making for all of us.) >>>> > > >>>> > > The spirit of this is really a one line change to the Rust policy: >>>> > > >>>> > > > It’s fine to use LLMs to answer questions, analyze, distill, >>>> refine, check, suggest, review. But not to create. >>>> > > >>>> > > ...becomes... >>>> > > >>>> > > It’s fine to use LLMs to answer questions, analyze, distill, >>>> refine, check, suggest, review. But not to decide. >>>> > > >>>> > > >>>> > > >>>> > > On Mon, Sep 28, 2026 at 12:54 PM Caleb Rackliffe < >>>> [email protected] <mailto:[email protected]>> wrote: >>>> > >> I finally read the Rust and Lucene policy docs in more detail... >>>> > >> >>>> > >> >>>> https://flagged.apple.com:443/proxy?t2=Dx1j3o8bN8&o=aHR0cHM6Ly9mb3JnZS5ydXN0LWxhbmcub3JnL3BvbGljaWVzL2xsbS11c2FnZS5odG1s&emid=10aadc4f-b52a-4662-9281-ce00061a230c&c=11 >>>> <https://flagged.apple.com/proxy?t2=Dx1j3o8bN8&o=aHR0cHM6Ly9mb3JnZS5ydXN0LWxhbmcub3JnL3BvbGljaWVzL2xsbS11c2FnZS5odG1s&emid=10aadc4f-b52a-4662-9281-ce00061a230c&c=11> >>>> < >>>> https://flagged.apple.com/proxy?t2=Dx1j3o8bN8&o=aHR0cHM6Ly9mb3JnZS5ydXN0LWxhbmcub3JnL3BvbGljaWVzL2xsbS11c2FnZS5odG1s&emid=10aadc4f-b52a-4662-9281-ce00061a230c&c=11 >>>> > >>>> > >> https://github.com/apache/lucene/blob/main/AI_POLICY.md >>>> > >> >>>> > >> I think I agree with a lot of what's written in both, and they >>>> overlap quite a lot, especially around communication (docs, issue comments, >>>> etc.) that should be primarily human-to-human. Everything useful in the >>>> Lucene policy is already included in the Rust policy though. If we could >>>> take the Rust policy, generalize away the Rust-specific things, simplify >>>> it, and remove the "Experimental" tag (and probably the "non-critical" >>>> qualifier) on the "LLM-created code changes intended of review" section, I >>>> think that's something a large majority of us would be able to live with. >>>> > >> >>>> > >> If we can get this right, it's simply clarifying the set of things >>>> contributors (including existing committers) can do to have the best chance >>>> at getting engagement from reviewers. >>>> > >> >>>> > >> I don't know how much appetite there is out there for a formal >>>> draft of this, and we already have 3-4 proposals, but I could attempt it if >>>> that would be useful... >>>> > >> >>>> > >> >>>> > >> On Mon, Sep 28, 2026 at 10:50 AM Štefan Miklošovič < >>>> [email protected] <mailto:[email protected]>> wrote: >>>> > >>> A clarification from my side, I asked "what is wrong with this" >>>> in my >>>> > >>> latest email: >>>> > >>> >>>> > >>> "For these reasons, it should be expected that the person >>>> producing >>>> > >>> the patch has already demonstrated their expertise and commitment >>>> by >>>> > >>> producing and maintaining similar patches without the use of AI" >>>> > >>> >>>> > >>> It is "almost fine", the part of "similar patches without the use >>>> of >>>> > >>> AI" should not be there. It should stop before that. >>>> > >>> >>>> > >>> Otherwise this is going to exclude people who have a decade of >>>> > >>> experience with Cassandra and contributed countless patches of >>>> various >>>> > >>> size and complexity while according to that exact wording, they >>>> would >>>> > >>> not be eligible to contribute an AI patch. That is silly. I think >>>> this >>>> > >>> is wrong. It does not matter how it was produced. What is >>>> important is >>>> > >>> established trust and if a patch is correct. What does even the >>>> size >>>> > >>> of a patch have in common with that? Expertise and commitment! Not >>>> > >>> "the series of patches this committer ever produced was not >>>> complex >>>> > >>> enough so we can't take that code in". >>>> > >>> >>>> > >>> Also, who is exactly going to measure that anyway? What are the >>>> > >>> _objective_ criteria who qualifies? Somebody might come and say >>>> "while >>>> > >>> based on my criteria, (because I do not like this person), I do >>>> not >>>> > >>> think that the patches of this person qualify, because ...". We >>>> need >>>> > >>> hard data on whether it can be merged or not, performance >>>> improvement, >>>> > >>> stability ... >>>> > >>> >>>> > >>> I think this particular wording would need to be refined further. >>>> > >>> >>>> > >>> On Mon, Sep 28, 2026 at 4:34 PM Štefan Miklošovič >>>> > >>> <[email protected] <mailto:[email protected]>> wrote: >>>> > >>> > >>>> > >>> > Right ... for that reason I don't think we should restrict >>>> anybody to >>>> > >>> > create a PR or anything like that, putting some artificial >>>> constraints >>>> > >>> > people will eventually bypass anyway. We don't have that under >>>> > >>> > control. What we have under control is the review part of that. >>>> A >>>> > >>> > patch not merged will not be released. The review itself is the >>>> > >>> > "gate". >>>> > >>> > >>>> > >>> > If a PR, even done by AI, is up to standards, has everything it >>>> should >>>> > >>> > have and it is technically correct, then I can not reject to >>>> merge >>>> > >>> > that only on the basis it was AI-generated. A patch like a >>>> patch. The >>>> > >>> > code speaks. The ultimate gate is if a patch is correct or not, >>>> not >>>> > >>> > how it was produced. >>>> > >>> > >>>> > >>> > Do I gravitate with my trust more towards established members >>>> of the >>>> > >>> > community? Definitely. The trust is earned over the years. >>>> Implicitly, >>>> > >>> > I am trusting a newcomer less. Sorry but not sorry. If somebody >>>> calls >>>> > >>> > this "gating", I don't think they see the nuances enough. Yeah, >>>> call >>>> > >>> > it a gate if you want ... >>>> > >>> > >>>> > >>> > That is why I agree with Benedict here, he said: >>>> > >>> > >>>> > >>> > "For these reasons, it should be expected that the person >>>> producing >>>> > >>> > the patch has already demonstrated their expertise and >>>> commitment by >>>> > >>> > producing and maintaining similar patches without the use of >>>> AI". >>>> > >>> > >>>> > >>> > What is wrong about this? >>>> > >>> > >>>> > >>> > Look at this contributor (1). This is an excellent example. 10 >>>> patches >>>> > >>> > in fast cadence three weeks ago. We never heard about this >>>> person >>>> > >>> > before nor after the patches were created. What about hitting a >>>> ML >>>> > >>> > saying "hey, guys, I have a set of patches which scratch my >>>> itches, >>>> > >>> > can you take a look, please?". I don't know ... just be a bit >>>> ... >>>> > >>> > human about all of this? The maintainers are people too. I am >>>> not >>>> > >>> > obliged to take in and cooperate with whoever comes by, dumps >>>> their >>>> > >>> > stuff and then they ... wait. Well, so wait. See where you got >>>> three >>>> > >>> > weeks after? Nowhere. >>>> > >>> > >>>> > >>> > Caleb put it nicely, we are "only" humans. >>>> > >>> > >>>> > >>> > (1) >>>> https://github.com/apache/cassandra/pulls?q=is%3Apr+state%3Aopen+author%3Acheeeee >>>> > >>> > >>>> > >>> > On Mon, Sep 28, 2026 at 2:51 PM Shailaja Koppu via dev >>>> > >>> > <[email protected] <mailto:[email protected]>> >>>> wrote: >>>> > >>> > > >>>> > >>> > > Hi Stefan, >>>> > >>> > > >>>> > >>> > > From personal side, I completely agree with you. I am giving >>>> potential options only to address concerns like - new contributors >>>> overwhelming the community with AI generated PRs just to show as add-on in >>>> their profile and vanish after that, or purely AI opened PRs without >>>> developer review or understanding. But the later can happen with anyone >>>> including committers due to workload/deadlines or misled by AI etc. Also, >>>> someone can copy a AI generated patch line by line skipping comments, which >>>> looks like a handwritten code. >>>> > >>> > > >>>> > >>> > > >>>> > >>> > > Thanks, >>>> > >>> > > Shailaja >>>> > >>> > > >>>> > >>> > > >>>> > >>> > > >>>> > >>> > > > On Sep 28, 2026, at 12:23 PM, Štefan Miklošovič < >>>> [email protected] <mailto:[email protected]>> wrote: >>>> > >>> > > > >>>> > >>> > > >> - Only Cassandra committers may submit AI-assisted PRs. >>>> This would mean new contributors first write and understand code without AI >>>> before becoming committers; or >>>> > >>> > > >> - Contributors may submit AI-assisted changes in a >>>> component/subcomponent only after they have submitted at least one non-AI >>>> PR in that component/subcomponent. >>>> > >>> > > > >>>> > >>> > > > I am not sure if I am missing something but can you all >>>> explain in >>>> > >>> > > > simple terms how is this actually enforceable in practice? >>>> > >>> > > > >>>> > >>> > > > "Only Cassandra committers may submit AI-assisted PRs" - >>>> there is no >>>> > >>> > > > restriction who can create a PR and how. It is not like we >>>> see that a >>>> > >>> > > > PR is created with heavy AI usage, then we check if a >>>> contributor is a >>>> > >>> > > > committer and when they are not we comment on that PR >>>> saying - "hold >>>> > >>> > > > your horses mate, we checked the list and you are not a >>>> committer, >>>> > >>> > > > sorry, we have to close this". >>>> > >>> > > > >>>> > >>> > > > If a PR is crafted "carefuly" then it might look like a >>>> completely >>>> > >>> > > > legitimate piece of work while it is still 100% prompted >>>> and the >>>> > >>> > > > author does not have a clue what they did. I mean ... how >>>> do you make >>>> > >>> > > > the difference between what is "real" and what is AI-driven >>>> 100%? I >>>> > >>> > > > think that even if we "guessed" which one is which, the >>>> possibility to >>>> > >>> > > > see this is being progressively erased as this tech is >>>> evolving and we >>>> > >>> > > > will eventually not have a clue. >>>> > >>> > > >>>> > >>>> > >>>> >>>>
