- Core code changes made by LLM may only be proposed by contributors with demonstrated expertise - Must have produced similar patches in size, scope and area unassisted and with minimal third-party guidance
I really don't like this one or its wording. Definitely too "the peasants are getting uppity lets build a wall". Lets not let a subjective thing like demonstrated expertise (who decides that?) be if it's ok or not. Hold the same standards for code quality and process for it all. I don't want this to be: only people on the storage team in Apple can use AI. Chris On Wed, Sep 23, 2026 at 12:16 AM <[email protected]> wrote: > I agree with Stefan and think this is both a reasonable and thoughtful > proposal. > > Here are some things I like about it: > > – It outlines areas where LLM usage is unambiguously useful to the > project’s developers and users. > – It defines a spectrum of recommendations and cautions. > – The only prohibited areas are extremely narrow and say nothing about > code at all. > > Some in this thread are responding as if this proposal seeks to prohibit > or sharply limit use of LLMs. In fact, it’s one of the most open and > welcoming I’ve seen for an OSS project of our size where many are adopting > policies that simply ban them entirely. I’ve re-appended the proposal below > my message as it seems to have been lost in threaded replies, and would > encourage folks to give it a second read. > > Some brief thoughts based on my own use of LLMs: > > – I find them fantastically useful for reviewing and identifying problems > that have slipped through review - primarily via Alex Petrov’s /deep-review > skill, which I have running in a VM in a loop executing over every new > commit in the project as of a few days ago. I will be posting a few > hand-authored Jira tickets based on findings that appear legitimate to me. > For now, the loop is posting them as issue drafts for my own review on my > personal fork which you can find here: > https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug > – They’re great for enabling use of model checkers and formal methods > where such work would have previously been prohibitively expensive, such as > Blake’s work on a TLA+ proof of aspects of Mutation Tracking and > Benedict/Fedor’s work on a machine-checkable proof of the Accord protocol > in Lean. > – They are stunning for allowing me to experiment with ideas that would > have otherwise been a summer internship’s scope of work. Some examples > include an io_uring prototype, exploring the impact of page-aligned > compressed chunk sizes, an API shim bridging the 3.x and 4.x Java Drivers, > and potential enhancements to Zstandard. > – And they shine when given grunt-work that is critical to the project but > a miserable labor for humans, such as triaging, reproducing, and > root-causing flaky tests, which David Capwell now has running in a loop to > help us improve CI stability in the project. > > I never thought I’d be so positive on what’s possible via language models > a year ago. At the same time, I also agree that they present challenges and > risks that can be managed through thoughtful discussion and policy. Some of > the concerns that I think are important to guard against include: > > *– Asymmetry of effort between author and reviewers:* As token-generating > machines, LLMs can generate diffs of extraordinary size very rapidly. > /deep-review is great for chewing through diffs and identifying defects. > But it should be used by the contributor themselves to identify issues – > not to replace the role of the reviewer with more electricity. The role of > the reviewers extends beyond identifying and highlighting defects. It > encompasses architecture, harmony with the existing codebase, thinking > ahead to future evolution of the project, and replicates context on the > project as new code is committed. These functions cannot be automated away. > *– Hesitancy of authors to engage manually with code they have generated:* > This is not specific to Cassandra, but it is a behavior that I have seen in > several “highly-electric” projects. There’s a bimodal tendency toward code > that is entirely generated or entirely human-authored - but it is rare for > someone to prepare an AI-authored patch to take an offramp and spend a > significant amount of time refining the work by hand in an IDE. This > hesitancy toward human participation in authorship of LLM-generated code is > very concerning to me. > *– Harmony with the existing codebase: *Due to the tunnel-vision of > context windows, LLMs are generally unaware of conventions and norms > present in codebases and very frequently reinvent concepts in a generation > turn to suit a goal without view of the project’s overall architecture. > This results in a profusion of messy and duplicated concepts that gradually > sprawl about a codebase. > > Again, none of these are grounds for prohibition of usage of language > models in developing the project. They’re just problems we need to bear in > mind and guard against – and I think the proposal is designed to do just > that. > > I’m thrilled by the potential of LLMs to improve Apache Cassandra and we > already see it happening through a vast number of issues that are being > reported and fixed. But there’s also danger in taking ATVs down a hiking > trail full of people. > > Regarding the prohibition on prose, I’ll simply say: I recently found > myself in a scenario where I found a Claude-authored document so > inscrutable that I piped it back into a model, directed it to rewrite it > in ASD-STE100, read it myself, and responded based on the summarization. As > a humanities grad, this is probably the worst language crime I have > committed. But it was in response to language that was itself so > idiosyncratic that it was unreadable to me in its original form. I hope > this never happens in the Apache Cassandra project. > > I’ll close with a quote from an excellent article written by Colin Breck, > an engineer who works on large-scale data systems: > https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/ > > Colin wrote: > > > I don’t want to live in a world where you use AI to summarize something > important into unreadable text, and then I use AI in an attempt to decipher > it. I want to hear you, imperfections and all. I want your interpretation > of aesthetics, beauty, quality, relationship, time. I want to know how you > feel. I want you to cut through and tell me what really matters. > > > Intentional writing will likely become more valuable. People who write, > and write to think, to think deeply and carefully, or to create, to share, > or to capture something important without explicitly expressing it will > continue to write and produce original work. The people who never were > writers will use AI to produce lots of text. > > I hope that our culture can remain one of intentional writing and > intentional engineering. I enjoy reading the voice of the author in > comments, code, and tickets in Cassandra – the different ways we use > language based on where we grew up and how we learned English, the > translated idioms from our various backgrounds, and terse comments that > recognize the difference between code whose function is obvious and what > warrants genuine exposition. When I read code in Cassandra, it’s a delight > to recognize the author based on their writing style before flipping on > `git annotate` to reveal the origin. > > I’d encourage folks to re-read the original proposal below. It is very > permissive. The guidance strikes me not just as reasonable, but genuinely > important to maintaining the health of the project. > > – Scott > > ===== > Encouraged: > - Reviewing and otherwise validating human-authored patches before > submission > - Debugging, diagnosing etc > > Permitted: > - Generating or modifying tests, scripts, tooling or any other non-user > facing changes > - Minor changes to human-authored patches that are carefully reviewed by > the author > > Restricted: > - Core code changes made by LLM may only be proposed by contributors with > demonstrated expertise > - Must have produced similar patches in size, scope and area unassisted > and with minimal third-party guidance > - Core code changes made by LLM require an additional reviewer > - LLM review is not a substitute for human review, and must be used only > to augment a complete and independent human understanding of the patch. > > Prohibited: > - All public prose must be human authored. This includes inline comments, > docs, posts to Jira etc. > > All LLM generated changes MUST be disclosed: > - Outlined to any reviewer; > - Summarised in the commit message; > - Large blocks or files must be individually marked with some agreed > message like "created by <some AI>" > ===== > > On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected]> wrote: > > On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected]> wrote: > >> >> Restricted: >> - Core code changes made by LLM may only be proposed by contributors with >> demonstrated expertise >> - Must have produced similar patches in size, scope and area >> unassisted and with minimal third-party guidance >> > > I am -1 on this. This sounds like gate keeping attempt. It narrowly limits > the pool to a few people on the project that have historically contributed > to certain parts of the codebase. This policy will prohibit skilled > software engineers with domain expertise from proposing LLM assisted > changes simply because they have not contributed to the project. This is > unrealistic and a net negative for the project to attract talent and grow > our community. > > >> - Core code changes made by LLM require an additional reviewer >> > > Can you be more precise what is this in addition to? How many total > reviewers do you expect and what is the purpose of additional reviewer? and > why? > > Taking a step back - what are you trying to solve here? > > Dinesh > > > >
