Hi Zhifeng, Thanks for reviving this KIP. I'm very supportive of the principle of being able to subdivide the brokers in a cluster for a variety of reasons. I have a few initial comments.
AS1: I do not think the pod should be exposed in the MetadataResponse and o.a.k.common.Node. Is it really useful to expose this concept to normal applications? The reason I am asking is that I think we might want to subdivide clusters in a way which is invisible to applications in the future. There is a multi-tenancy KIP and I imagine that we would want to be able to scope tenants to a subset of brokers. AS2: It seems a bit inelegant having the canary configuration in the controller configs and also in the arguments to kafka-reassign-partitions.sh. It would be much preferable if the tool could discover the canary information from the cluster to ensure consistency. AS3: I wonder whether using "pod" when it is so widely used in the context of Kubernetes is sensible. I suggest "cell" instead, but this is just my personal opinion. Thanks, Andrew On 2026/08/06 07:02:27 Chen Zhifeng wrote: > Title: [DISCUSS] KIP-1095 Kafka Canary Isolation > > Hi Everyone, > > Apologies for the long silence. Reviving KIP-1095 > <https://cwiki.apache.org/confluence/spaces/KAFKA/pages/323488210/KIP-1095+Kafka+Canary+Isolation> > after a substantial revision. The original thread is at here > <https://lists.apache.org/thread/n7mprq43fh39hsgj48bzfbtbho3l1cpy>; and > thanks Divij for the questions there. > > *What changed since initial discussion* > - Scope narrowed to broker-side placement - producer/consumer-side > improvements as future work > - Cleaned up zookeeper related changes > - Protocol changes reframed as tagged fields - no client upgrade required > > *Answering the earlier questions* > > Why can’t we achieve the objective without making any change at all? For > example, you can designate a few brokers as your “canary brokers” where > your custom "canary partitions" are situated. During rolling deployment you > can choose to deploy changes to these brokers at the beginning. If the > health of your canary partitions is good, you can continue ahead with the > rest of deployment. > A: The suggested deployment process has been used at Uber for years, and > this KIP is developed on top of it. What operating it showed is that > designating brokers alone does not give you isolation between the > designated brokers and the rest. > A partition's replicas span brokers, so unless placement guarantees that > the entire replica set lands inside the canary pod, a canary partition > still has followers on non-canary brokers, and vice versa. Without that > isolation, impact leaks from one broker to every other broker it is > connected to — a leader running new code can propagate bad data to > followers, and a degraded follower can slow down a leader that was never > upgraded. The blast radius of the deployment therefore grows from a small, > known percentage of requests to an unknown and unbounded portion. > Detection suffers for the same reason. When one partition degrades, > producers of keyless records simply route around it to healthy partitions, > so a partial failure may not be detectable until the rollout is nearly > complete. > > > What do you think about having a separate canary cluster where you deploy > code first before deploying to production cluster. The canary cluster could > receive a small portion of "shadow" production traffic or have it's own > synthetic traffic. > A: Shadowing, or capture/replay, is another approach to safe deployment, > with a different trade-off: > - capture/replay pays for redundancy in order to keep impact away from > production entirely, which requires extra storage, compute, and engineering > cost; > - canary isolation instead bounds the blast radius of a bad deployment to > a pre-calculated portion of traffic, at lower engineering cost and with no > extra hardware. > One important difference is that shadow traffic does not reproduce > production behaviour exactly, so some regressions will always slip through > to production. Canary traffic is production traffic, so it inherits > production behaviour by construction. The two are complementary, and which > one fits depends on the operator; this KIP aims to make the second option > available in Kafka itself. > > > Would controller broker be part of canary brokers or not? How will we > test code regression in controller? Similarly how will we test code > regression in transaction coordinator and consumer coordinator? > A: controllers are not covered by this KIP. In KRaft the controller is a > separate role with its own quorum and rolling procedure, and the > canary-partition concept does not map onto it; catching controller > regressions needs a different mechanism, which I would rather not fold into > this proposal. > > The coordinators are a different case. Both the group and transaction > coordinators are partition leaders of __consumer_offsets and > __transaction_state, so the mechanism in this KIP reaches them by applying > the same placement rules to those internal topics. I have left that out of > the initial scope deliberately, but I am happy to discuss it as follow-up > work if there is interest. > > *Production status*: The idea has been applied at Uber for around 2 years. > with 1/32 partitions being canary partitions, blast radius of bad kafka > deployment has been contained with-in ~3% of production. > > *Draft implementation*: https://github.com/apache/kafka/pull/23095. > > *JIRA*: https://issues.apache.org/jira/browse/KAFKA-20897. > > Thanks, > Zhifeng >
