Hi, Sounds great, thanks Gyula!
I created the following JIRA [1] to wrap up what we discussed here. I plan to address it as part of the next 1.17.0 operator release. Thank you all for the feedback! Best regards, Dennis [1] https://issues.apache.org/jira/browse/FLINK-40455 On Fri, Aug 21, 2026 at 2:37 PM Gyula Fóra <[email protected]> wrote: > Hey Dennis! > > I think we all agree here, let's just start with some required > refactoring, flow improvements over the current model to make it more > extensible. > The we can move to hardening the logic and adding more capabilities > such as RocksDB memory tuning on the reworked flow. > > Later we can discuss splitting it into it's own module once this has > stabilized and we still feel the need. > > Since this is not API breaking or any fundamentally new large feature > I think we don't need a separate FLIP but maybe an umbrella jira to > track the work more easily. > > Cheers > Gyula > > On Wed, Aug 5, 2026 at 4:37 PM Dennis-Mircea Ciupitu > <[email protected]> wrote: > > > > Thanks Max! Glad we agree on the potential of the Autotuner and on making > > changes to it. I think we also agree on moving in small steps, so let me > > make the first increment concrete. To remove any ambiguity about a big > > rework: I am dropping the module split and any public pluggable API from > > the proposal entirely. What remains is one focused change: decoupling the > > tuning decision from the scaling decision, within the existing autoscaler > > module, keeping the tuning logic self-contained and extensible > internally. > > > > That decoupling is what fixes the biggest gap raised in this thread: jobs > > that keep parallelism fixed, by intent or by external constraint, never > get > > any tuning today, and jobs where the autoscaler finds nothing to change > are > > never re-tuned either. > > > > Step 1 in concrete terms: > > * Tuning evaluation gets its own decision step next to the scaling one, > > inside the same reconcile loop, with its own trigger interval, cooldown > > and minimum-improvement guard. > > * Coordination stays as tight as today, by construction: metrics are > > collected once, the scaling decision runs first, and the tuning > decision > > runs after it, seeing the scaling outcome. When scaling triggers a > > restart anyway, pending tuning is applied in the same cycle. Otherwise > > tuning fires only on its own guards. Scaling keeps precedence > throughout. > > * The tuning logic moves behind a small internal boundary so it can be > > tested and stabilized on its own. Today it can only be observed through > > the scaling path. > > * No new Maven module, no separate coordinator component, no public API, > > no change to the tuning algorithm. Default behavior stays exactly as > > today unless the new trigger is enabled. > > > > On RocksDB memory tuning as the starting point: I would make it the first > > tuner on top of step 1, even in the same release, rather than the first > > step itself. Any tuner added inside the current shape only computes when > > the autoscaler decides to rescale, and RocksDB tuning targets exactly the > > jobs Francis described, large state, where rescaling is avoided or > > impractical. With a scaling-only trigger it would almost never run for > > exactly the jobs it is meant to help. Right after it I would place CPU > > tuning, since > > CPU is the other half of the container cost and a CPU-bound job at fixed > > parallelism gets no help from memory tuning at all. The JobManager > > counterparts would follow. > > > > On restart modes: agreed, and I think we are describing the same work > from > > two ends. The mechanisms are already there (in-place rescale vs the full > > upgrade cycle). Making them accessible from the Autotuner code requires > an > > Autotuner code path to call them from, and today the only tuning > call-site > > lives inside ScalingExecutor, tied to a scaling decision. Step 1 creates > > exactly that decision point, including the option of holding a > > recommendation until the next restart that happens anyway. > > > > Does this match what you both have in mind as the first increment? If > so, I > > will trim the proposal down to exactly this scope, and I am fine either > > keeping it as a small FLIP or tracking it as JIRAs under FLINK-34538, > > whichever you prefer. The individual tuners (RocksDB, CPU, JM) would then > > be filed as separate follow-up JIRAs, each with its own justification and > > rollout. > > > > Best regards, > > Dennis > > > > > > On Tue, Aug 4, 2026 at 11:26 AM Maximilian Michels <[email protected]> > wrote: > > > > > Thanks Dennis for starting the conversation! There's a lot of > > > potential in the Autotuner. I very much support making changes to it. > > > > > > I would prefer if we moved in small steps instead of breaking apart > > > everything. For example, why don't we start by adding RocksDB memory > > > tuning? > > > > > > Concerning the different ways Autotuner / Autoscaler apply their > > > changes (full redeploy vs in-place), I think we can add explicit > > > restart modes. They are already there, we just need to make them > > > accessible from the Autotuner code. > > > > > > Cheers, > > > Max > > > > > > On Fri, Jul 31, 2026 at 11:33 AM Dennis-Mircea Ciupitu > > > <[email protected]> wrote: > > > > > > > > Hi Gyula, Francis, all, > > > > > > > > Thank you for the thoughts on this topic, they are much appreciated. > To > > > > restate the proposal before I answer the two main concerns: tuning > would > > > > get its own decision loop and trigger, while staying coordinated with > > > > scaling. > > > > > > > > 1. Decoupling autotuning from autoscaling > > > > > > > > Today having autotuning running only when autoscaler decides to > change > > > > parallelism has 2 consequences: > > > > • Jobs that intentionally keep parallelism fixed never get any > tuning at > > > > all. As Francis said, good integration between the two is desirable, > but > > > > the autotuner is also valuable on its own for pipelines that cannot > or do > > > > not want to move parallelism: pod quotas per namespace, alignment > with > > > > external systems (partition counts, sink constraints), or simply jobs > > > where > > > > rescaling is expensive due to large state. > > > > • The tuning logic is hard to test, stabilize and evaluate, because > we > > > can > > > > only observe it through the scaling path. That limits how much > confidence > > > > we can build in it, and how much value we can extract from it. > > > > > > > > Two concrete scenarios that decoupling unlocks: > > > > • Infra cost savings: a pipeline that runs continuously at its > minimum > > > > parallelism can still have its container memory brought down to what > it > > > > actually needs when traffic is low. > > > > • Infra cost optimization: when parallelism is capped by external > > > > constraints, vertical adjustment is the only remaining lever, and we > can > > > > right-size the TM replicas running at that fixed parallelism instead > of > > > > leaving them over-provisioned. Letting tuning also raise resources > for > > > such > > > > jobs, within the limits the user configured, is an extension I would > like > > > > this FLIP to cover, since today the initially configured memory acts > as > > > an > > > > upper bound. > > > > > > > > There are of course mixed cases in between, but the real value of the > > > > decoupling is exactly this: two clearly separated levels of > optimization, > > > > parallelism-level and resource-level. I understand the concern about > > > > complexity. My expectation is that the responsibilities become more > > > > explicit this way, with each level having its own trigger conditions, > > > > stabilization/cooldown and metrics. Concretely, the decoupled tuner > would > > > > keep using the same metric collection infrastructure as the > autoscaler. > > > The > > > > evaluation, decision, and its lifecycle are separated, so this is > not a > > > > parallel subsystem. > > > > > > > > On restarts, I think the two cases differ in an important way: > in-place > > > > rescaling never redeploys the cluster, while a memory change always > does, > > > > today as well, since a JVM cannot resize its heap. The two should > indeed > > > be > > > > combined exactly where that pays off: when the scaling itself cannot > be > > > > done in-place and requires a restart anyway, the pending tuning > decision > > > > should be applied in the same cycle, so we pay for one restart > instead of > > > > two. But that is an opportunistic optimization rather than a reason > to > > > keep > > > > the two tied together. When scaling happens in-place there is no > restart > > > to > > > > share, and applying a tuning decision at that moment forces a full > > > upgrade > > > > cycle that would not otherwise have happened, which cancels out the > > > benefit > > > > of in-place rescaling. So tuning needs its own trigger, cooldown and > > > > minimum-improvement threshold in any case, plus the option to hold a > > > > pending recommendation until the next restart that is going to happen > > > > anyway. The current behaviour and recommendation-only mode stay as > the > > > > default. > > > > I would also spell out the coordination rules in the FLIP: scaling > > > > decisions take precedence over tuning ones, and tuning runs on a > longer > > > > interval with its own cooldown, so it cannot chase the effects of a > > > recent > > > > rescale. That also covers the feedback between the two, since memory > > > > changes influence restart times, which in turn feed back into scaling > > > > decisions. > > > > > > > > 2. The pluggable interface > > > > > > > > The current autotuning implementation is not behind any interface > and the > > > > code surface is fairly small. The intent of introducing one is not to > > > > expose a public extension point to end users at this stage. It is an > > > > internal seam, and I would rather justify it with the use cases it is > > > meant > > > > to serve: > > > > • TaskManager CPU: memory is only one half of the container cost, > and a > > > job > > > > that is CPU-bound at a fixed parallelism cannot be helped by memory > > > tuning > > > > at all. > > > > • JobManager resources: JM sizing is static today and tends to be > > > > over-provisioned, which adds up on platforms running many small jobs. > > > > > > > > Each of these is a separate decision, with its own metrics and its > own > > > risk > > > > profile. A common seam lets us add and test them one at a time > without > > > > touching the scaling logic, and roll them out independently. > > > > > > > > There is also a timing argument, and it is the reason I wanted to > raise > > > > this early rather than after the fact. The tuning implementation is > still > > > > small and has not been extended much yet, which makes this the > cheapest > > > > moment to restructure it. Once JM tuning, CPU tuning and further > resource > > > > optimizations are built on top of the current shape, the same > migration > > > > becomes considerably more invasive, and it is exactly the kind of > change > > > > nobody wants to make on a feature that users already depend on. > > > > > > > > I would keep the interface internal and outside the public API > surface to > > > > begin with, and only revisit making it user-facing once the baseline > is > > > > proven. This is in line with what was preferred in the FLIP-543 > > > discussion: > > > > adding extension points rather than replacing core components. > > > > > > > > To make this easier to review, I would sequence the work as follows: > > > > (a) extract tuning from the scaling path, with its own trigger and > > > guards, > > > > behaviour unchanged by default; > > > > (b) introduce the internal interface in the same step, since it is > > > cheapest > > > > to add exactly while the surface is this small; > > > > (c) add the individual tuners incrementally afterwards, each with > its own > > > > justification and rollout. > > > > > > > > That way the structural change lands once, on a small surface, and > > > > everything built on top of it can be discussed and merged separately. > > > > > > > > Looking forward to your feedback. > > > > > > > > Best regards, > > > > Dennis > > > > > > > > În joi, 30 iul. 2026 la 13:09 Francis Altomare < > > > > [email protected]> a scris: > > > > > > > > > Hey all, > > > > > Thanks for the discussion here. I wanted to provide my experience > here > > > as > > > > > well since I think our applications could benefit from a decoupled > > > > > autocaler and autotuner. > > > > > > > > > > For applications with very large keyed state stores, autoscaling > as it > > > > > currently works is not practical. Even when using ForSt and async > state > > > > > APIs. While upscaling operations are fast, downscaling operations > can > > > be > > > > > incredibly slow due to what I believe is a blocking repartition of > the > > > > > underlying state store. > > > > > > > > > > Since autotuning shouldn’t impact job parallelism, it should be a > safe > > > > > scaling operation regardless of the state store size. Decoupling > this > > > > > feature could give larger stateful applications a safer way to use > it. > > > > > > > > > > Thanks, > > > > > Francis > > > > > > > > > > > On Jul 30, 2026, at 10:48, Gyula Fóra <[email protected]> > wrote: > > > > > > > > > > > > Hey Dennis! > > > > > > > > > > > > I think the autotuner (memory scaling) is not nearly as well > > > developed > > > > > and > > > > > > stable as the autoscaler (parallelism scaling) is. > > > > > > At the moment, my personal opinion is that we should leave it as > is > > > and > > > > > > further develop the functionality and streamline the overall > scaling > > > > > > approach. I have a few reasons why I think this: > > > > > > > > > > > > 1. Although we have introduced two names (autoscaler/autotuner) > these > > > > > > functionalities are very much related. They both work based on > > > metrics > > > > > and > > > > > > execute reconfigurations to adjust the job performance. When and > how > > > > > these > > > > > > are executed have a huge impact on downtime and reliability so a > > > tight > > > > > > integration is essential. Breaking this up into two modules and a > > > > > > coordinator may just make this more complicated and harder to > > > maintain. > > > > > > > > > > > > 2. I am really against adding any pluggable interfaces to the > > > existing > > > > > > logic until the current ones are well developed. I have expressed > > > this > > > > > > opinion on the autscaler logic as well a few times that I don't > think > > > > > > pluggable interfaces on somewhat experimental/non-battle tasted > > > features > > > > > > serve the project or the community. We need to develop a solid > core > > > > > feature > > > > > > set , get adoption then consider pluggability later. > > > > > > > > > > > > You also called out a few other smaller changes like > improvements to > > > > > > configs etc but those can be done incrementally even in the > current > > > > > setup, > > > > > > also they feel somewhat unnecessary at this point (kind of going > > > back to > > > > > > the question of stable featureset, adoption, trust) etc. > > > > > > > > > > > > Cheers > > > > > > Gyula > > > > > > > > > > > > On Mon, Jul 6, 2026 at 10:52 AM Dennis-Mircea Ciupitu < > > > > > > [email protected]> wrote: > > > > > > > > > > > >> Hi all, > > > > > >> > > > > > >> I’d like to start a discussion on FLIP-XXX: Flink Kubernetes > > > Operator > > > > > >> Autotuning Redesign [1]. > > > > > >> > > > > > >> Flink Autotuning stands out as one of the most powerful > components > > > the > > > > > >> operator can offer, and as the place where the real cost > > > optimizations > > > > > of a > > > > > >> Flink platform are achieved. Autoscaling decides how many > resources > > > a > > > > > job > > > > > >> runs on, while autotuning decides how efficiently every > provisioned > > > > > >> resource is used. Fully developed, it continuously right-sizes > > > memory, > > > > > task > > > > > >> slots, JobManager resources, and CPU for every managed job, > without > > > > > >> changing job behavior and without requiring users to be experts > in > > > > > Flink's > > > > > >> memory model. > > > > > >> > > > > > >> Today, however, autotuning ships as a sub-feature of Flink > > > Autoscaling. > > > > > In > > > > > >> practice this means it cannot run when the autoscaler is > disabled, > > > fresh > > > > > >> recommendations are only computed when a scaling decision is > > > actually > > > > > made, > > > > > >> and there is no seam for adding new tuning types. The tuning > roadmap > > > > > under > > > > > >> FLINK-34538 [2] reflects this: JobManager memory (FLINK-34539) > and > > > task > > > > > >> slots (FLINK-34540) tuning have been open and unassigned for > about > > > two > > > > > >> years, and RocksDB-aware or CPU tuning have no tickets at all. > > > > > >> > > > > > >> The FLIP proposes promoting autotuning to a first-class > optimizer, > > > > > >> following the precedent of FLIP-334 [3], which decoupled the > > > autoscaler > > > > > >> from the operator. In short: > > > > > >> > > > > > >> * A new flink-autotuner module as a peer of flink-autoscaler, > plus a > > > > > small > > > > > >> shared flink-optimizer-common base module. > > > > > >> * A pluggable Tuner interface, so new tuning types (JobManager > > > memory, > > > > > task > > > > > >> slots, RocksDB, CPU) can be added without touching the > autoscaler & > > > > > >> autotuner. > > > > > >> * A target-aware ConfigChanges type, so the realizer can > dispatch > > > > > >> JobManager, TaskManager, Flink config, and Kubernetes changes > > > without > > > > > >> key-prefix heuristics. > > > > > >> * An operator-side coordinator that runs both optimizers on a > single > > > > > metric > > > > > >> snapshot per reconcile. The autotuner runs on every reconcile, > and > > > > > >> enablement of the two is fully orthogonal. > > > > > >> * A dedicated job.autotuner.* config namespace, with the > existing > > > > > >> job.autoscaler.memory.tuning.* keys kept as deprecated aliases, > > > plus a > > > > > new > > > > > >> floor-ratio option so the spec can act as a configurable lower > > > bound for > > > > > >> memory recommendations. > > > > > >> > > > > > >> The current TaskManager memory tuning algorithm itself is ported > > > as-is. > > > > > >> This FLIP is about the architecture, not the tuning math. > > > > > >> > > > > > >> Looking forward to your thoughts and suggestions! > > > > > >> > > > > > >> Best regards, > > > > > >> Dennis > > > > > >> > > > > > >> [1] > > > > > >> > > > > > >> > > > > > > > > > https://docs.google.com/document/d/1JuNa-_7NwwL-GjqWHm8nnUuFpdNtGuLJj4Zfnl8m_O0/edit?usp=sharing > > > > > >> [2] https://issues.apache.org/jira/browse/FLINK-34538 > > > > > >> [3] > > > > > >> > > > > > >> > > > > > > > > > https://cwiki.apache.org/confluence/spaces/FLINK/pages/263424711/FLIP-334+Decoupling+autoscaler+and+kubernetes+and+support+the+Standalone+Autoscaler > > > > > >> > > > > > > > > > > > > > >
