[ 
https://issues.apache.org/jira/browse/FLINK-40789?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18124945#comment-18124945
 ] 

Martijn Visser commented on FLINK-40789:
----------------------------------------

Two more on master. Push run 2026-10-05 (test_ci core, JDK 17), 
testRecordInProgressRescale:
https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79792&view=logs&j=0da23115-68bb-5dcd-192c-bd4c8adebde1

{code}
[ERROR] 
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.testRecordInProgressRescale
 -- Time elapsed: 0.543 s <<< FAILURE!
org.opentest4j.AssertionFailedError:
expected: null
 but was: NO_RESOURCES_OR_PARALLELISMS_CHANGE
        at 
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.lambda$testRecordInProgressRescale$11(RescaleTimelineITCase.java:386)
{code}

Same cooldown race as the master run in my last comment: the cooldown ended 1 
ms after the UPDATE_REQUIREMENT rescale was created, so it was already 
terminated when the test took the snapshot:

{code}
13:48:33,303 ... Rescale [] - Updated rescale is: Rescale{... 
triggerCause=UPDATE_REQUIREMENT, terminalState=null, terminatedReason=null}
13:48:33,304 ... DefaultStateTransitionManager [] - Transitioning from Cooldown 
to Idling, job a59d5fc22ec6d46289dd2f3064b4e17f.
13:48:33,318 ... Current rescale Rescale{... triggerCause=UPDATE_REQUIREMENT, 
terminalState=IGNORED, terminatedReason=NO_RESOURCES_OR_PARALLELISMS_CHANGE} is 
null or terminated, so the update action is ignored.
{code}

Nightly 2026-10-08 (test_cron_hadoop343 core, JDK 17), 
testRecordRescaleForNewResourcesRequirements:
https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79930&view=logs&j=5af3fd07-293c-5932-0967-2a3127c0f840

{code}
[ERROR] 
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.testRecordRescaleForNewResourcesRequirements
 -- Time elapsed: 0.494 s <<< FAILURE!
org.opentest4j.AssertionFailedError:
expected: SUCCEEDED
 but was: null
        at 
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCaseBase.assertTerminalRelatedFields(RescaleTimelineITCaseBase.java:151)
        at 
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.lambda$testRecordRescaleForNewResourcesRequirements$1(RescaleTimelineITCase.java:139)
{code}

Here the job was again first scheduled at parallelism 2 instead of 4, so the 
requirement update to 1..2 did not change the parallelism and the 
UPDATE_REQUIREMENT rescale was still open when the test took the snapshot. It 
ended with JOB_FINISHED after the unblock:

{code}
02:03:07,330 ... Rescale{... rescaleAttemptId=1}, ... 
preRescaleParallelism=null, postRescaleParallelism=2 ... 
triggerCause=INITIAL_SCHEDULE, terminalState=COMPLETED, 
terminatedReason=SUCCEEDED}
02:03:07,417 ... Rescale [] - Updated rescale is: Rescale{... 
triggerCause=UPDATE_REQUIREMENT, terminalState=null, terminatedReason=null}
02:03:07,426 ... Rescale{... rescaleAttemptId=1}, ... preRescaleParallelism=2, 
postRescaleParallelism=null ... triggerCause=UPDATE_REQUIREMENT, 
terminalState=IGNORED, terminatedReason=JOB_FINISHED}
{code}

That is four of the 101 master and release-2.3 runs in the last seven days.

> RescaleTimelineITCase.testRecordRescaleForNewAvailableResource records three 
> rescales instead of two
> ----------------------------------------------------------------------------------------------------
>
>                 Key: FLINK-40789
>                 URL: https://issues.apache.org/jira/browse/FLINK-40789
>             Project: Flink
>          Issue Type: Bug
>          Components: Runtime / Coordination
>    Affects Versions: 2.3.1, 2.4.0
>            Reporter: Martijn Visser
>            Assignee: Naveen Joseph
>            Priority: Critical
>              Labels: test-stability
>
> https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79368&view=logs&j=0da23115-68bb-5dcd-192c-bd4c8adebde1
>  (master push 2026-09-23, test_ci core, JDK 17)
> {code}
> [ERROR] 
> org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.testRecordRescaleForNewAvailableResource
>  -- Time elapsed: 0.722 s <<< FAILURE!
> java.lang.AssertionError:
> Expected size: 2 but was: 3 in:
> [Rescale{... rescaleAttemptId=3, ... preRescaleParallelism=4, 
> postRescaleParallelism=6 ... triggerCause=NEW_RESOURCE_AVAILABLE, 
> terminalState=COMPLETED ...},
>     Rescale{... rescaleAttemptId=2, ... preRescaleParallelism=2, 
> postRescaleParallelism=4 ... triggerCause=NEW_RESOURCE_AVAILABLE, 
> terminalState=COMPLETED ...},
>     Rescale{... rescaleAttemptId=1, ... preRescaleParallelism=null, 
> postRescaleParallelism=2 ... triggerCause=INITIAL_SCHEDULE, 
> terminalState=COMPLETED ...}]
>         at 
> org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.lambda$testRecordRescaleForNewAvailableResource$2(RescaleTimelineITCase.java:170)
>         at 
> org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.runAdaptedParameterizedAssertion(RescaleTimelineITCase.java:397)
>         at 
> org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.testRecordRescaleForNewAvailableResource(RescaleTimelineITCase.java:164)
> {code}
>                                                                               
>                                                                               
>                    The test expects the initial schedule at parallelism 4 and 
> one upscale to 6 after the third TaskManager starts. Here the job was first 
> scheduled at 2 and went 2 -> 4 -> 6. The other RescaleTimelineITCase methods 
> had similar races fixed in FLINK-40010, FLINK-40067 and FLINK-40076; this 
> method has no ticket.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to