- Core code changes made by LLM may only be proposed by contributors with
demonstrated expertise
   - Must have produced similar patches in size, scope and area unassisted
and with minimal third-party guidance

I really don't like this one or its wording. Definitely too "the peasants
are getting uppity lets build a wall". Lets not let a subjective thing like
demonstrated expertise (who decides that?) be if it's ok or not. Hold the
same standards for code quality and process for it all. I don't want this
to be: only people on the storage team in Apple can use AI.

Chris

On Wed, Sep 23, 2026 at 12:16 AM <[email protected]> wrote:

> I agree with Stefan and think this is both a reasonable and thoughtful
> proposal.
>
> Here are some things I like about it:
>
> – It outlines areas where LLM usage is unambiguously useful to the
> project’s developers and users.
> – It defines a spectrum of recommendations and cautions.
> – The only prohibited areas are extremely narrow and say nothing about
> code at all.
>
> Some in this thread are responding as if this proposal seeks to prohibit
> or sharply limit use of LLMs. In fact, it’s one of the most open and
> welcoming I’ve seen for an OSS project of our size where many are adopting
> policies that simply ban them entirely. I’ve re-appended the proposal below
> my message as it seems to have been lost in threaded replies, and would
> encourage folks to give it a second read.
>
> Some brief thoughts based on my own use of LLMs:
>
> – I find them fantastically useful for reviewing and identifying problems
> that have slipped through review - primarily via Alex Petrov’s /deep-review
> skill, which I have running in a VM in a loop executing over every new
> commit in the project as of a few days ago. I will be posting a few
> hand-authored Jira tickets based on findings that appear legitimate to me.
> For now, the loop is posting them as issue drafts for my own review on my
> personal fork which you can find here:
> https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug
> – They’re great for enabling use of model checkers and formal methods
> where such work would have previously been prohibitively expensive, such as
> Blake’s work on a TLA+ proof of aspects of Mutation Tracking and
> Benedict/Fedor’s work on a machine-checkable proof of the Accord protocol
> in Lean.
> – They are stunning for allowing me to experiment with ideas that would
> have otherwise been a summer internship’s scope of work. Some examples
> include an io_uring prototype, exploring the impact of page-aligned
> compressed chunk sizes, an API shim bridging the 3.x and 4.x Java Drivers,
> and potential enhancements to Zstandard.
> – And they shine when given grunt-work that is critical to the project but
> a miserable labor for humans, such as triaging, reproducing, and
> root-causing flaky tests, which David Capwell now has running in a loop to
> help us improve CI stability in the project.
>
> I never thought I’d be so positive on what’s possible via language models
> a year ago. At the same time, I also agree that they present challenges and
> risks that can be managed through thoughtful discussion and policy. Some of
> the concerns that I think are important to guard against include:
>
> *– Asymmetry of effort between author and reviewers:* As token-generating
> machines, LLMs can generate diffs of extraordinary size very rapidly.
> /deep-review is great for chewing through diffs and identifying defects.
> But it should be used by the contributor themselves to identify issues –
> not to replace the role of the reviewer with more electricity. The role of
> the reviewers extends beyond identifying and highlighting defects. It
> encompasses architecture, harmony with the existing codebase, thinking
> ahead to future evolution of the project, and replicates context on the
> project as new code is committed. These functions cannot be automated away.
> *– Hesitancy of authors to engage manually with code they have generated:*
> This is not specific to Cassandra, but it is a behavior that I have seen in
> several “highly-electric” projects. There’s a bimodal tendency toward code
> that is entirely generated or entirely human-authored - but it is rare for
> someone to prepare an AI-authored patch to take an offramp and spend a
> significant amount of time refining the work by hand in an IDE. This
> hesitancy toward human participation in authorship of LLM-generated code is
> very concerning to me.
> *– Harmony with the existing codebase: *Due to the tunnel-vision of
> context windows, LLMs are generally unaware of conventions and norms
> present in codebases and very frequently reinvent concepts in a generation
> turn to suit a goal without view of the project’s overall architecture.
> This results in a profusion of messy and duplicated concepts that gradually
> sprawl about a codebase.
>
> Again, none of these are grounds for prohibition of usage of language
> models in developing the project. They’re just problems we need to bear in
> mind and guard against – and I think the proposal is designed to do just
> that.
>
> I’m thrilled by the potential of LLMs to improve Apache Cassandra and we
> already see it happening through a vast number of issues that are being
> reported and fixed. But there’s also danger in taking ATVs down a hiking
> trail full of people.
>
> Regarding the prohibition on prose, I’ll simply say: I recently found
> myself in a scenario where I found a Claude-authored document so
> inscrutable that I piped it back into a model, directed it to rewrite it
> in ASD-STE100, read it myself, and responded based on the summarization. As
> a humanities grad, this is probably the worst language crime I have
> committed. But it was in response to language that was itself so
> idiosyncratic that it was unreadable to me in its original form. I hope
> this never happens in the Apache Cassandra project.
>
> I’ll close with a quote from an excellent article written by Colin Breck,
> an engineer who works on large-scale data systems:
> https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/
>
> Colin wrote:
>
> > I don’t want to live in a world where you use AI to summarize something
> important into unreadable text, and then I use AI in an attempt to decipher
> it. I want to hear you, imperfections and all. I want your interpretation
> of aesthetics, beauty, quality, relationship, time. I want to know how you
> feel. I want you to cut through and tell me what really matters.
>
> > Intentional writing will likely become more valuable. People who write,
> and write to think, to think deeply and carefully, or to create, to share,
> or to capture something important without explicitly expressing it will
> continue to write and produce original work. The people who never were
> writers will use AI to produce lots of text.
>
> I hope that our culture can remain one of intentional writing and
> intentional engineering. I enjoy reading the voice of the author in
> comments, code, and tickets in Cassandra – the different ways we use
> language based on where we grew up and how we learned English, the
> translated idioms from our various backgrounds, and terse comments that
> recognize the difference between code whose function is obvious and what
> warrants genuine exposition. When I read code in Cassandra, it’s a delight
> to recognize the author based on their writing style before flipping on
> `git annotate` to reveal the origin.
>
> I’d encourage folks to re-read the original proposal below. It is very
> permissive. The guidance strikes me not just as reasonable, but genuinely
> important to maintaining the health of the project.
>
> – Scott
>
> =====
> Encouraged:
> - Reviewing and otherwise validating human-authored patches before
> submission
> - Debugging, diagnosing etc
>
> Permitted:
> - Generating or modifying tests, scripts, tooling or any other non-user
> facing changes
> - Minor changes to human-authored patches that are carefully reviewed by
> the author
>
> Restricted:
> - Core code changes made by LLM may only be proposed by contributors with
> demonstrated expertise
>    - Must have produced similar patches in size, scope and area unassisted
> and with minimal third-party guidance
> - Core code changes made by LLM require an additional reviewer
> - LLM review is not a substitute for human review, and must be used only
> to augment a complete and independent human understanding of the patch.
>
> Prohibited:
> - All public prose must be human authored. This includes inline comments,
> docs, posts to Jira etc.
>
> All LLM generated changes MUST be disclosed:
> - Outlined to any reviewer;
> - Summarised in the commit message;
> - Large blocks or files must be individually marked with some agreed
> message like "created by <some AI>"
> =====
>
> On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected]> wrote:
>
> On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected]> wrote:
>
>>
>> Restricted:
>> - Core code changes made by LLM may only be proposed by contributors with
>> demonstrated expertise
>>     - Must have produced similar patches in size, scope and area
>> unassisted and with minimal third-party guidance
>>
>
> I am -1 on this. This sounds like gate keeping attempt. It narrowly limits
> the pool to a few people on the project that have historically contributed
> to certain parts of the codebase. This policy will prohibit skilled
> software engineers with domain expertise from proposing LLM assisted
> changes simply because they have not contributed to the project. This is
> unrealistic and a net negative for the project to attract talent and grow
> our community.
>
>
>> - Core code changes made by LLM require an additional reviewer
>>
>
> Can you be more precise what is this in addition to? How many total
> reviewers do you expect and what is the purpose of additional reviewer? and
> why?
>
> Taking a step back - what are you trying to solve here?
>
> Dinesh
>
>
>
>

Reply via email to