I was waiting for this moment to hit our project and I'm glad we're here. I
made it difficult to contribute. I had hoped that this new era of
more diverse thoughts and ideas. This policy proposal is the exact opposite
of what we need. We have been sitting on a Cassandra 6 release alpha for
months. We need to accelerate and embrace new ways of being or be left
behind. As I read that policy, my first and gut level reactions:
- It comes across as elitist and class protectionism. Committer should not
be special but this proposal makes that designation even more sacred.
understand it" That's some SQLite vibes right there.
Sadly, i think this policy change would also exclude a lot of comitters.
We aren't alone in this moment. The Linux project just went through
this. You can find the thread with a simple Google, but similar hard
bases. Human or Human using AI. “You are expected to understand and to be
able to defend everything you submit.” Love that.
In the larger picture, I'll restate. I'm worried for our project. In late
like jet fuel. Here's some examples of new projects being hyper fueled by
AI coding tools.
Apache Iggy - Complete rust replacement of kafka. Crazy fast velocity
Bun - Rust re-write of itself from Zig.
startup, but it was him alone using a ton of local AI coding agents. He
even implemented Accord. Yeah...
The cracks are already starting to show. There is a black market economy of
Cassandra patches happening now. Not going to name names or call people
the Cassandra project. Why? I'll use myself as an example. I fixed a nasty
bug I ran into with TCM a few weeks ago. Wrote the tests. It passes CI and
lives in my personal branch. I'm sitting here really wondering if I want to
go through the ritual humiliation of being roasted for using AI to fix it.
Me. I am worried about contrinuting code the Cassandra. What the hell does
month. I would love to donate that to the Cassandra project but I wouldn't
if it essentially killed any progress.
everything you submit.” approach the Linux project has adopted.
- Loosen up the contributor process and our worry on trunk. Let 1000
flowers bloom and bring it in.
adopt what other projects have done and provide more pluggability. Let new
ideas have an easy place to connect.
We are at a fork in the road. What are we going to do? And then I have to
I’m not necessarily opposed to having a policy, but so far we have some
specific proposals addressing a problem statement that’s very nebulous.
What is the community failing to do on its own that we’re trying to correct
with policy? What outcomes are we trying to create or prevent? Having some
examples and specific problems to discuss would help focus the conversation.
On Wed, Sep 23, 2026, at 4:34 AM, Shailaja Koppu via dev wrote:
Benedict,
Thanks for clarifying. My concern still remains. This criteria would be
difficult to define and apply consistently. What counts as “similar” scope
or area, “mostly correct,” or sufficiently independent work? More
importantly, how do we prevent such vague criteria from creating an
informal hierarchy where some contributors work is routinely accepted while
others is routinely rejected?
If the intent is to limit AI-assisted code changes to Cassandra
contributors, or to contributors who have previously worked in that
component without AI, that would at least be clear and enforceable.
On Sep 23, 2026, at 12:01 PM, Benedict Elliott Smith <
Core code changes
Chris: Do you object to the first or second line you quote? Because the
first line is effectively motivation for the second line, and can be
removed (or more clearly combined). If it’s the second line, then I do not
think this is an unreasonable expectation, and we can get into a proper
debate about it.
Shailaja, since you only snipped the first sentence, your concerns might
also be mostly answered by this clarification? “Minimal third-party
guidance” implies you have some concerns about the second line, but all of
our policies have some ambiguity because legalese is even worse. I don’t
think the ambiguity here would be challenging to navigate though we can
certainly refine it. This specific snippet is meant to convey an
expectation that a contributor has autonomously produced patches of similar
scope that were mostly correct, so that they have demonstrated the level of
understanding necessary to guide another party to a successful patch (i.e.
an LLM in this case).
On 2026/09/23 10:54:16 Benedict Elliott Smith wrote:
Thanks everyone for your input so far. I’ll respond in brief to the
main themes, in (mostly) separate emails so they can each have their own
debate chain.
Should we have a policy (Blake/Josh*/Jon/Dinesh)
I think we would all agree that LLMs represent the biggest change to
this community (and software more generally) since its inception, and we
all now have enough experience with the technology to have formed opinions
about how it is best managed. We also evidently have not all arrived at the
same conclusions. In this situation, it would be an abdication of our
responsibilities as a management committee to not agree *some* policy.
I intend to conduct straw polls as the discussion evolves, so if you
prefer an alternative policy - or modifications to this policy - I would
encourage you to make those alternative proposals.
*Veto/Consensus (Josh)
It was fair to call out my poor use of language on this topic, so let
me rephrase a little. The community is built on consensus, and work should
not be merged when there are outstanding concerns to address. The explicit
-1 should only be used rarely, because the prior expectation should prevent
it ever being needed. I (and others) have outstanding concerns on LLM
generated work that can only be addressed through this process right here,
so to merge such work while maintaining the community’s consensus we must
agree some policy.
On 2026/09/23 09:58:27 Shailaja Koppu via dev wrote:
I am strongly -1 on this
- Core code changes made by LLM may only be proposed by contributors
with demonstrated expertise
That creates a new, subjective privileged class of contributors and
turns a tool choice into an eligibility test. Who decides whether expertise
has been “demonstrated,” what counts as “minimal third-party guidance,” and
how could those judgments be applied consistently or fairly?
Apache already has a better model, anyone may contribute, trust and
additional repository privileges are earned transparently over time. The
ASF describes its communities as flat, and says that newcomer ideas have as
much input as those from original creators. We should not add a separate,
informal hierarchy in which certain people may use common development tools
while others may not.
On Sep 23, 2026, at 6:33 AM, Chris Lohfink <[email protected]>
wrote:
- Core code changes made by LLM may only be proposed by contributors
with demonstrated expertise
- Must have produced similar patches in size, scope and area
unassisted and with minimal third-party guidance
I really don't like this one or its wording. Definitely too "the
peasants are getting uppity lets build a wall". Lets not let a subjective
thing like demonstrated expertise (who decides that?) be if it's ok or not.
Hold the same standards for code quality and process for it all. I don't
want this to be: only people on the storage team in Apple can use AI.
Chris
I agree with Stefan and think this is both a reasonable and
thoughtful proposal.
Here are some things I like about it:
– It outlines areas where LLM usage is unambiguously useful to the
project’s developers and users.
– It defines a spectrum of recommendations and cautions.
– The only prohibited areas are extremely narrow and say nothing
about code at all.
Some in this thread are responding as if this proposal seeks to
prohibit or sharply limit use of LLMs. In fact, it’s one of the most open
and welcoming I’ve seen for an OSS project of our size where many are
adopting policies that simply ban them entirely. I’ve re-appended the
proposal below my message as it seems to have been lost in threaded
replies, and would encourage folks to give it a second read.
Some brief thoughts based on my own use of LLMs:
– I find them fantastically useful for reviewing and identifying
problems that have slipped through review - primarily via Alex Petrov’s
/deep-review skill, which I have running in a VM in a loop executing over
every new commit in the project as of a few days ago. I will be posting a
few hand-authored Jira tickets based on findings that appear legitimate to
me. For now, the loop is posting them as issue drafts for my own review on
my personal fork which you can find here:
– They’re great for enabling use of model checkers and formal
methods where such work would have previously been prohibitively expensive,
such as Blake’s work on a TLA+ proof of aspects of Mutation Tracking and
Benedict/Fedor’s work on a machine-checkable proof of the Accord protocol
in Lean.
– They are stunning for allowing me to experiment with ideas that
would have otherwise been a summer internship’s scope of work. Some
examples include an io_uring prototype, exploring the impact of
page-aligned compressed chunk sizes, an API shim bridging the 3.x and 4.x
Java Drivers, and potential enhancements to Zstandard.
– And they shine when given grunt-work that is critical to the
project but a miserable labor for humans, such as triaging, reproducing,
and root-causing flaky tests, which David Capwell now has running in a loop
to help us improve CI stability in the project.
I never thought I’d be so positive on what’s possible via language
models a year ago. At the same time, I also agree that they present
challenges and risks that can be managed through thoughtful discussion and
policy. Some of the concerns that I think are important to guard against
include:
– Asymmetry of effort between author and reviewers: As
token-generating machines, LLMs can generate diffs of extraordinary size
very rapidly. /deep-review is great for chewing through diffs and
identifying defects. But it should be used by the contributor themselves to
identify issues – not to replace the role of the reviewer with more
electricity. The role of the reviewers extends beyond identifying and
highlighting defects. It encompasses architecture, harmony with the
existing codebase, thinking ahead to future evolution of the project, and
replicates context on the project as new code is committed. These functions
cannot be automated away.
– Hesitancy of authors to engage manually with code they have
generated: This is not specific to Cassandra, but it is a behavior that I
have seen in several “highly-electric” projects. There’s a bimodal tendency
toward code that is entirely generated or entirely human-authored - but it
is rare for someone to prepare an AI-authored patch to take an offramp and
spend a significant amount of time refining the work by hand in an IDE.
This hesitancy toward human participation in authorship of LLM-generated
code is very concerning to me.
– Harmony with the existing codebase: Due to the tunnel-vision of
context windows, LLMs are generally unaware of conventions and norms
present in codebases and very frequently reinvent concepts in a generation
turn to suit a goal without view of the project’s overall architecture.
This results in a profusion of messy and duplicated concepts that gradually
sprawl about a codebase.
Again, none of these are grounds for prohibition of usage of
language models in developing the project. They’re just problems we need to
bear in mind and guard against – and I think the proposal is designed to do
just that.
I’m thrilled by the potential of LLMs to improve Apache Cassandra
and we already see it happening through a vast number of issues that are
being reported and fixed. But there’s also danger in taking ATVs down a
hiking trail full of people.
Regarding the prohibition on prose, I’ll simply say: I recently
found myself in a scenario where I found a Claude-authored document so
inscrutable that I piped it back into a model, directed it to rewrite it in
ASD-STE100, read it myself, and responded based on the summarization. As a
humanities grad, this is probably the worst language crime I have
committed. But it was in response to language that was itself so
idiosyncratic that it was unreadable to me in its original form. I hope
this never happens in the Apache Cassandra project.
I’ll close with a quote from an excellent article written by Colin
Breck, an engineer who works on large-scale data systems:
Colin wrote:
I don’t want to live in a world where you use AI to summarize
something important into unreadable text, and then I use AI in an attempt
to decipher it. I want to hear you, imperfections and all. I want your
interpretation of aesthetics, beauty, quality, relationship, time. I want
to know how you feel. I want you to cut through and tell me what really
matters.
Intentional writing will likely become more valuable. People who
write, and write to think, to think deeply and carefully, or to create, to
share, or to capture something important without explicitly expressing it
will continue to write and produce original work. The people who never were
writers will use AI to produce lots of text.
I hope that our culture can remain one of intentional writing and
intentional engineering. I enjoy reading the voice of the author in
comments, code, and tickets in Cassandra – the different ways we use
language based on where we grew up and how we learned English, the
translated idioms from our various backgrounds, and terse comments that
recognize the difference between code whose function is obvious and what
warrants genuine exposition. When I read code in Cassandra, it’s a delight
to recognize the author based on their writing style before flipping on
`git annotate` to reveal the origin.
I’d encourage folks to re-read the original proposal below. It is
very permissive. The guidance strikes me not just as reasonable, but
genuinely important to maintaining the health of the project.
– Scott
=====
Encouraged:
- Reviewing and otherwise validating human-authored patches before
submission
- Debugging, diagnosing etc
Permitted:
- Generating or modifying tests, scripts, tooling or any other
non-user facing changes
- Minor changes to human-authored patches that are carefully
reviewed by the author
Restricted:
- Core code changes made by LLM may only be proposed by contributors
with demonstrated expertise
- Must have produced similar patches in size, scope and area
unassisted and with minimal third-party guidance
- Core code changes made by LLM require an additional reviewer
- LLM review is not a substitute for human review, and must be used
only to augment a complete and independent human understanding of the patch.
Prohibited:
- All public prose must be human authored. This includes inline
comments, docs, posts to Jira etc.
All LLM generated changes MUST be disclosed:
- Outlined to any reviewer;
- Summarised in the commit message;
- Large blocks or files must be individually marked with some agreed
message like "created by <some AI>"
=====
On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected]
On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected]
Restricted:
- Core code changes made by LLM may only be proposed by
contributors with demonstrated expertise
- Must have produced similar patches in size, scope and area
unassisted and with minimal third-party guidance
I am -1 on this. This sounds like gate keeping attempt. It narrowly
limits the pool to a few people on the project that have historically
contributed to certain parts of the codebase. This policy will prohibit
skilled software engineers with domain expertise from proposing LLM
assisted changes simply because they have not contributed to the project.
This is unrealistic and a net negative for the project to attract talent
and grow our community.
- Core code changes made by LLM require an additional reviewer
Can you be more precise what is this in addition to? How many total
reviewers do you expect and what is the purpose of additional reviewer? and
why?
Taking a step back - what are you trying to solve here?
Dinesh