Hi Zhifeng,
Thanks for reviving this KIP. I'm very supportive of the principle of being 
able to subdivide the brokers in a cluster for a variety of reasons. I have a 
few initial comments.

AS1: I do not think the pod should be exposed in the MetadataResponse and 
o.a.k.common.Node. Is it really useful to expose this concept to normal 
applications? The reason I am asking is that I think we might want to subdivide 
clusters in a way which is invisible to applications in the future. There is a 
multi-tenancy KIP and I imagine that we would want to be able to scope tenants 
to a subset of brokers.

AS2: It seems a bit inelegant having the canary configuration in the controller 
configs and also in the arguments to kafka-reassign-partitions.sh. It would be 
much preferable if the tool could discover the canary information from the 
cluster to ensure consistency.

AS3: I wonder whether using "pod" when it is so widely used in the context of 
Kubernetes is sensible. I suggest "cell" instead, but this is just my personal 
opinion. 

Thanks,
Andrew

On 2026/08/06 07:02:27 Chen Zhifeng wrote:
> Title: [DISCUSS] KIP-1095 Kafka Canary Isolation
> 
> Hi Everyone,
> 
> Apologies for the long silence. Reviving KIP-1095
> <https://cwiki.apache.org/confluence/spaces/KAFKA/pages/323488210/KIP-1095+Kafka+Canary+Isolation>
> after a substantial revision. The original thread is at here
> <https://lists.apache.org/thread/n7mprq43fh39hsgj48bzfbtbho3l1cpy>; and
> thanks Divij for the questions there.
> 
> *What changed since initial discussion*
>   - Scope narrowed to broker-side placement - producer/consumer-side
> improvements as future work
>   - Cleaned up zookeeper related changes
>   - Protocol changes reframed as tagged fields - no client upgrade required
> 
> *Answering the earlier questions*
> > Why can’t we achieve the objective without making any change at all? For
> example, you can designate a few brokers as your “canary brokers” where
> your custom "canary partitions" are situated. During rolling deployment you
> can choose to deploy changes to these brokers at the beginning. If the
> health of your canary partitions is good, you can continue ahead with the
> rest of deployment.
> A: The suggested deployment process has been used at Uber for years, and
> this KIP is developed on top of it. What operating it showed is that
> designating brokers alone does not give you isolation between the
> designated brokers and the rest.
> A partition's replicas span brokers, so unless placement guarantees that
> the entire replica set lands inside the canary pod, a canary partition
> still has followers on non-canary brokers, and vice versa. Without that
> isolation, impact leaks from one broker to every other broker it is
> connected to — a leader running new code can propagate bad data to
> followers, and a degraded follower can slow down a leader that was never
> upgraded. The blast radius of the deployment therefore grows from a small,
> known percentage of requests to an unknown and unbounded portion.
> Detection suffers for the same reason. When one partition degrades,
> producers of keyless records simply route around it to healthy partitions,
> so a partial failure may not be detectable until the rollout is nearly
> complete.
> 
> > What do you think about having a separate canary cluster where you deploy
> code first before deploying to production cluster. The canary cluster could
> receive a small portion of "shadow" production traffic or have it's own
> synthetic traffic.
> A: Shadowing, or capture/replay, is another approach to safe deployment,
> with a different trade-off:
>  - capture/replay pays for redundancy in order to keep impact away from
> production entirely, which requires extra storage, compute, and engineering
> cost;
>  - canary isolation instead bounds the blast radius of a bad deployment to
> a pre-calculated portion of traffic, at lower engineering cost and with no
> extra hardware.
> One important difference is that shadow traffic does not reproduce
> production behaviour exactly, so some regressions will always slip through
> to production. Canary traffic is production traffic, so it inherits
> production behaviour by construction. The two are complementary, and which
> one fits depends on the operator; this KIP aims to make the second option
> available in Kafka itself.
> 
> > Would controller broker be part of canary brokers or not? How will we
> test code regression in controller? Similarly how will we test code
> regression in transaction coordinator and consumer coordinator?
> A:  controllers are not covered by this KIP. In KRaft the controller is a
> separate role with its own quorum and rolling procedure, and the
> canary-partition concept does not map onto it; catching controller
> regressions needs a different mechanism, which I would rather not fold into
> this proposal.
> 
> The coordinators are a different case. Both the group and transaction
> coordinators are partition leaders of __consumer_offsets and
> __transaction_state, so the mechanism in this KIP reaches them by applying
> the same placement rules to those internal topics. I have left that out of
> the initial scope deliberately, but I am happy to discuss it as follow-up
> work if there is interest.
> 
> *Production status*: The idea has been applied at Uber for around 2 years.
> with 1/32 partitions being canary partitions, blast radius of bad kafka
> deployment has been contained with-in ~3% of production.
> 
> *Draft implementation*: https://github.com/apache/kafka/pull/23095.
> 
> *JIRA*: https://issues.apache.org/jira/browse/KAFKA-20897.
> 
> Thanks,
> Zhifeng
> 

Reply via email to