[ 
https://issues.apache.org/jira/browse/FLINK-40455?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Dennis-Mircea Ciupitu updated FLINK-40455:
------------------------------------------
    Description: 
h1. Background

Autotuning is currently applied only as part of autoscaling decisions, so new 
tuning recommendations are computed only when the autoscaler decides to 
rescale, and not at all when the autoscaler is disabled or doesn't take a 
scaling decision. This in practice limits the following concrete scenarios:
 * Jobs that keep parallelism fixed by design cannot benefit from any tuning.
 * Jobs that were scaled down to the lowest parallelism possible cannot be 
further tuned such that the resource allocation can be optimized down even more 
over long periods of low traffic.
 * Jobs that were scaled up to the highest parallelism possible cannot be 
further tuned such that the resource allocation can be optimized up even more 
over long periods of low traffic.

With the above limitations, autotuning logic is unable to further realize cost 
and infrastructure optimizations. Therefore, those limitations bring the need 
to have the autotuning decoupled from autoscaling logic, with coordination in 
between.

As part of the [dev mailing list 
discussion|https://lists.apache.org/thread/vk6j8793rx2y78vhh1yvfjf4qntpfkm9] we 
agreed to take the direction of decoupling autotuning from autoscaling logic 
within the autoscaler module.
h1. Goal

The main purpose of this task is to cover the following:
 * Autotuning model is refactored to a more extensible, self-contained model.
 * Autotuning logic is applied either within the autoscaling logic or as a 
separate step after autoscaling, depending on the case, under a full 
coordination mechanism.
 * Autoscaler metric collection will be shared by both autoscaling and the 
autotuning.
 * Autotuning is getting its own trigger interval, cooldown, and 
minimum-improvement guard.

By implementing the above refactorings, autotuning will have a stronger 
baseline on top of which multiple autotuners can be built. This will also open 
the door to further enhancements and capabilities, making it an even more 
self-contained, stable, and reliable feature, capable of delivering real cost 
and infrastructure optimizations.
h1. Out of scope

The following items/improvements will stay on hold until the autotuning gets 
further adoption and gets stabilized enough:
 * Move the autotuner into its own module.
 * Store the autotuning's `configOverrides` state into a dedicated autotuner 
ConfigMap.

  was:
h1. Background

Autotuning is currently applied only as part of autoscaling decisions, so new 
tuning recommendations are computed only when the autoscaler decides to 
rescale, and not at all when the autoscaler is disabled or doesn't take a 
scaling decision. This in practice limits the following concrete scenarios:
 * Jobs that keep parallelism fixed by design cannot benefit from any tuning.
 * Jobs that were scaled down to the lowest parallelism possible cannot be 
further tuned such that the resource allocation can be optimized down even more 
over long periods of low traffic.
 * Jobs that were scaled up to the highest parallelism possible cannot be 
further tuned such that the resource allocation can be optimized up even more 
over long periods of low traffic.

With the above limitations, autotuning logic is unable to further realize cost 
and infrastructure optimizations. Therefore, those limitations bring the need 
to have the autotuning decoupled from autoscaling logic, with coordination in 
between.

As part of the [dev mailing list 
discussion|https://lists.apache.org/thread/vk6j8793rx2y78vhh1yvfjf4qntpfkm9] we 
agreed to take the direction of decoupling autotuning from autoscaling logic 
within the autoscaler module.
h1. Goal

The main purpose of this task is to cover the following:
 * Autotuning model is refactored to a more extensible, self-contained model.
 * Autotuning logic is applied either within the autoscaling logic or as a 
separate step after autoscaling, depending on the case, under a full 
coordination mechanism.
 * Autoscaler metric collection will be shared by both autoscaling and the 
autotuning.
 * Autotuning is getting its own trigger interval, cooldown, and 
minimum-improvement guard.

By implementing the above refactoring, autotuning will have a stronger baseline 
on top of which multiple autotuners can be built. This will also open the door 
to further enhancements and capabilities, making it an even more 
self-contained, stable, and reliable feature, capable of delivering real cost 
and infrastructure optimizations.
h1. Out of scope

The following items/improvements will stay on hold until the autotuning gets 
further adoption and gets stabilized enough:
 * Move the autotuner into its own module.
 * Store the autotuning's `configOverrides` state into a dedicated autotuner 
ConfigMap.


> Decoupling Autotuning from the Autoscaling logic
> ------------------------------------------------
>
>                 Key: FLINK-40455
>                 URL: https://issues.apache.org/jira/browse/FLINK-40455
>             Project: Flink
>          Issue Type: Improvement
>          Components: Autoscaler, Kubernetes Operator
>            Reporter: Dennis-Mircea Ciupitu
>            Priority: Major
>
> h1. Background
> Autotuning is currently applied only as part of autoscaling decisions, so new 
> tuning recommendations are computed only when the autoscaler decides to 
> rescale, and not at all when the autoscaler is disabled or doesn't take a 
> scaling decision. This in practice limits the following concrete scenarios:
>  * Jobs that keep parallelism fixed by design cannot benefit from any tuning.
>  * Jobs that were scaled down to the lowest parallelism possible cannot be 
> further tuned such that the resource allocation can be optimized down even 
> more over long periods of low traffic.
>  * Jobs that were scaled up to the highest parallelism possible cannot be 
> further tuned such that the resource allocation can be optimized up even more 
> over long periods of low traffic.
> With the above limitations, autotuning logic is unable to further realize 
> cost and infrastructure optimizations. Therefore, those limitations bring the 
> need to have the autotuning decoupled from autoscaling logic, with 
> coordination in between.
> As part of the [dev mailing list 
> discussion|https://lists.apache.org/thread/vk6j8793rx2y78vhh1yvfjf4qntpfkm9] 
> we agreed to take the direction of decoupling autotuning from autoscaling 
> logic within the autoscaler module.
> h1. Goal
> The main purpose of this task is to cover the following:
>  * Autotuning model is refactored to a more extensible, self-contained model.
>  * Autotuning logic is applied either within the autoscaling logic or as a 
> separate step after autoscaling, depending on the case, under a full 
> coordination mechanism.
>  * Autoscaler metric collection will be shared by both autoscaling and the 
> autotuning.
>  * Autotuning is getting its own trigger interval, cooldown, and 
> minimum-improvement guard.
> By implementing the above refactorings, autotuning will have a stronger 
> baseline on top of which multiple autotuners can be built. This will also 
> open the door to further enhancements and capabilities, making it an even 
> more self-contained, stable, and reliable feature, capable of delivering real 
> cost and infrastructure optimizations.
> h1. Out of scope
> The following items/improvements will stay on hold until the autotuning gets 
> further adoption and gets stabilized enough:
>  * Move the autotuner into its own module.
>  * Store the autotuning's `configOverrides` state into a dedicated autotuner 
> ConfigMap.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to