+1 on the separate-repo placement — the Flink kubernetes-operator
precedent is the right shape here. A dedicated repo gives the operator
its own release cadence and issue tracker, which is important for a
component that will iterate on K8s API versions independently of Fluss
core.

Looking forward to the VOTE thread.

Best regards,
Zhe Wang


Lorenzo Affetti via dev <[email protected]> 于2026年5月5日周二 17:07写道:

> Hello community!
>
> I would like to start a discussion about FIP-41: Fluss Kubernetes Operator.
> Here is the motivation:
>
> Fluss 0.8 introduced a Helm chart
> <https://github.com/apache/fluss/issues/779> that simplifies the initial
> deployment of a Fluss cluster on Kubernetes by packaging manifests,
> configuration, and dependencies into a versioned release. While this is a
> good foundation, a Helm chart is fundamentally a one-shot templating tool.
> It has no awareness of Fluss's runtime state and cannot react to
> operational events such as pod failures, rolling upgrades, or scale-in
> operations that risk data loss.
>
> Running Fluss in production on Kubernetes today requires users to manually
> coordinate:
>
>    - Safe rolling restarts of TabletServers, ensuring tablet leadership is
>    migrated before each pod terminates (Fluss 0.8 introduced the graceful
>    shutdown
>    <
> https://fluss.apache.org/docs/maintenance/operations/graceful-shutdown/>
> primitive
>    that makes this possible, but does not orchestrate it across pods)
>    - Version upgrades that must sequence CoordinatorServer and TabletServer
>    updates correctly
>    - Scale-in operations where a TabletServer must be drained of tablets
>    before its pod is terminated
>    - Recovery from partial failures (e.g., PVC reattachment, pod identity
>    preservation across restarts)
>    - Leveraging Fluss 0.8's dynamic configuration updates
>    <https://fluss.apache.org/docs/maintenance/operations/updating-configs/>
> for
>    keys that do not require restart, instead of triggering a rolling
> restart
>    for every config change
>
> This gap between deployment and operations is best addressed by a
> *Kubernetes
> Operator* — a controller that continuously reconciles the desired state of
> a FlussCluster resource against the actual state of the cluster, and
> executes Fluss-aware transitions safely.
>
> The Flink and Spark ecosystems have established this pattern successfully
> with the Apache Flink Kubernetes Operator
> <https://github.com/apache/flink-kubernetes-operator> and the Apache Spark
> Kubernetes Operator <https://github.com/apache/spark-kubernetes-operator>.
> Fluss should follow suit.
>
> --
> Lorenzo Affetti
> Team Leader of Stream Storage
> [email protected]
> www.ververica.com
> ------------------------------
>
> <https://www.ververica.com/>
> Ververica GmbH | Herzogspitalstrasse 24 | 80331 München | Germany
>
> Follow us:
> <https://www.linkedin.com/company/ververica/posts/?feedView=all>
> <https://www.youtube.com/@ververica>
> <https://open.spotify.com/show/2XME9h8iBOyr6YupqM99ir?si=87b064644add42a1
> >Available
> on:  <https://aws.amazon.com/marketplace/pp/prodview-luvmqd6leha4i>
> <
> https://marketplace.microsoft.com/en-us/product/saas/ververica.vvc_managed?tab=Overview
> >
>
> Pflichtangaben/Mandatory Information
> <https://www.ververica.com/mandatory-information>
>

Reply via email to