I agree with Stefan and think this is both a reasonable and thoughtful proposal.

Here are some things I like about it:

– It outlines areas where LLM usage is unambiguously useful to the project’s 
developers and users.
– It defines a spectrum of recommendations and cautions.
– The only prohibited areas are extremely narrow and say nothing about code at 
all.

Some in this thread are responding as if this proposal seeks to prohibit or 
sharply limit use of LLMs. In fact, it’s one of the most open and welcoming 
I’ve seen for an OSS project of our size where many are adopting policies that 
simply ban them entirely. I’ve re-appended the proposal below my message as it 
seems to have been lost in threaded replies, and would encourage folks to give 
it a second read.

Some brief thoughts based on my own use of LLMs:

– I find them fantastically useful for reviewing and identifying problems that 
have slipped through review - primarily via Alex Petrov’s /deep-review skill, 
which I have running in a VM in a loop executing over every new commit in the 
project as of a few days ago. I will be posting a few hand-authored Jira 
tickets based on findings that appear legitimate to me. For now, the loop is 
posting them as issue drafts for my own review on my personal fork which you 
can find here: 
https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug
– They’re great for enabling use of model checkers and formal methods where 
such work would have previously been prohibitively expensive, such as Blake’s 
work on a TLA+ proof of aspects of Mutation Tracking and Benedict/Fedor’s work 
on a machine-checkable proof of the Accord protocol in Lean.
– They are stunning for allowing me to experiment with ideas that would have 
otherwise been a summer internship’s scope of work. Some examples include an 
io_uring prototype, exploring the impact of page-aligned compressed chunk 
sizes, an API shim bridging the 3.x and 4.x Java Drivers, and potential 
enhancements to Zstandard.
– And they shine when given grunt-work that is critical to the project but a 
miserable labor for humans, such as triaging, reproducing, and root-causing 
flaky tests, which David Capwell now has running in a loop to help us improve 
CI stability in the project.

I never thought I’d be so positive on what’s possible via language models a 
year ago. At the same time, I also agree that they present challenges and risks 
that can be managed through thoughtful discussion and policy. Some of the 
concerns that I think are important to guard against include:

– Asymmetry of effort between author and reviewers: As token-generating 
machines, LLMs can generate diffs of extraordinary size very rapidly. 
/deep-review is great for chewing through diffs and identifying defects. But it 
should be used by the contributor themselves to identify issues – not to 
replace the role of the reviewer with more electricity. The role of the 
reviewers extends beyond identifying and highlighting defects. It encompasses 
architecture, harmony with the existing codebase, thinking ahead to future 
evolution of the project, and replicates context on the project as new code is 
committed. These functions cannot be automated away.
– Hesitancy of authors to engage manually with code they have generated: This 
is not specific to Cassandra, but it is a behavior that I have seen in several 
“highly-electric” projects. There’s a bimodal tendency toward code that is 
entirely generated or entirely human-authored - but it is rare for someone to 
prepare an AI-authored patch to take an offramp and spend a significant amount 
of time refining the work by hand in an IDE. This hesitancy toward human 
participation in authorship of LLM-generated code is very concerning to me.
– Harmony with the existing codebase: Due to the tunnel-vision of context 
windows, LLMs are generally unaware of conventions and norms present in 
codebases and very frequently reinvent concepts in a generation turn to suit a 
goal without view of the project’s overall architecture. This results in a 
profusion of messy and duplicated concepts that gradually sprawl about a 
codebase.

Again, none of these are grounds for prohibition of usage of language models in 
developing the project. They’re just problems we need to bear in mind and guard 
against – and I think the proposal is designed to do just that.

I’m thrilled by the potential of LLMs to improve Apache Cassandra and we 
already see it happening through a vast number of issues that are being 
reported and fixed. But there’s also danger in taking ATVs down a hiking trail 
full of people.

Regarding the prohibition on prose, I’ll simply say: I recently found myself in 
a scenario where I found a Claude-authored document so inscrutable that I piped 
it back into a model, directed it to rewrite it in ASD-STE100, read it myself, 
and responded based on the summarization. As a humanities grad, this is 
probably the worst language crime I have committed. But it was in response to 
language that was itself so idiosyncratic that it was unreadable to me in its 
original form. I hope this never happens in the Apache Cassandra project.

I’ll close with a quote from an excellent article written by Colin Breck, an 
engineer who works on large-scale data systems: 
https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/

Colin wrote:

> I don’t want to live in a world where you use AI to summarize something 
> important into unreadable text, and then I use AI in an attempt to decipher 
> it. I want to hear you, imperfections and all. I want your interpretation of 
> aesthetics, beauty, quality, relationship, time. I want to know how you feel. 
> I want you to cut through and tell me what really matters.

> Intentional writing will likely become more valuable. People who write, and 
> write to think, to think deeply and carefully, or to create, to share, or to 
> capture something important without explicitly expressing it will continue to 
> write and produce original work. The people who never were writers will use 
> AI to produce lots of text.

I hope that our culture can remain one of intentional writing and intentional 
engineering. I enjoy reading the voice of the author in comments, code, and 
tickets in Cassandra – the different ways we use language based on where we 
grew up and how we learned English, the translated idioms from our various 
backgrounds, and terse comments that recognize the difference between code 
whose function is obvious and what warrants genuine exposition. When I read 
code in Cassandra, it’s a delight to recognize the author based on their 
writing style before flipping on `git annotate` to reveal the origin.

I’d encourage folks to re-read the original proposal below. It is very 
permissive. The guidance strikes me not just as reasonable, but genuinely 
important to maintaining the health of the project.

– Scott

=====
Encouraged:
- Reviewing and otherwise validating human-authored patches before submission
- Debugging, diagnosing etc

Permitted:
- Generating or modifying tests, scripts, tooling or any other non-user facing 
changes
- Minor changes to human-authored patches that are carefully reviewed by the 
author

Restricted:
- Core code changes made by LLM may only be proposed by contributors with 
demonstrated expertise
   - Must have produced similar patches in size, scope and area unassisted and 
with minimal third-party guidance
- Core code changes made by LLM require an additional reviewer
- LLM review is not a substitute for human review, and must be used only to 
augment a complete and independent human understanding of the patch.

Prohibited:
- All public prose must be human authored. This includes inline comments, docs, 
posts to Jira etc.

All LLM generated changes MUST be disclosed:
- Outlined to any reviewer;
- Summarised in the commit message;
- Large blocks or files must be individually marked with some agreed message 
like "created by <some AI>"
=====

> On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected]> wrote:
> 
> On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected] 
> <mailto:[email protected]>> wrote:
>> 
>> Restricted:
>> - Core code changes made by LLM may only be proposed by contributors with 
>> demonstrated expertise
>>     - Must have produced similar patches in size, scope and area unassisted 
>> and with minimal third-party guidance
> 
> I am -1 on this. This sounds like gate keeping attempt. It narrowly limits 
> the pool to a few people on the project that have historically contributed to 
> certain parts of the codebase. This policy will prohibit skilled software 
> engineers with domain expertise from proposing LLM assisted changes simply 
> because they have not contributed to the project. This is unrealistic and a 
> net negative for the project to attract talent and grow our community.
>  
>> - Core code changes made by LLM require an additional reviewer
> 
> Can you be more precise what is this in addition to? How many total reviewers 
> do you expect and what is the purpose of additional reviewer? and why?
> 
> Taking a step back - what are you trying to solve here?
> 
> Dinesh
>  

Reply via email to