[
https://issues.apache.org/jira/browse/FLINK-40789?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18124945#comment-18124945
]
Martijn Visser commented on FLINK-40789:
----------------------------------------
Two more on master. Push run 2026-10-05 (test_ci core, JDK 17),
testRecordInProgressRescale:
https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79792&view=logs&j=0da23115-68bb-5dcd-192c-bd4c8adebde1
{code}
[ERROR]
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.testRecordInProgressRescale
-- Time elapsed: 0.543 s <<< FAILURE!
org.opentest4j.AssertionFailedError:
expected: null
but was: NO_RESOURCES_OR_PARALLELISMS_CHANGE
at
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.lambda$testRecordInProgressRescale$11(RescaleTimelineITCase.java:386)
{code}
Same cooldown race as the master run in my last comment: the cooldown ended 1
ms after the UPDATE_REQUIREMENT rescale was created, so it was already
terminated when the test took the snapshot:
{code}
13:48:33,303 ... Rescale [] - Updated rescale is: Rescale{...
triggerCause=UPDATE_REQUIREMENT, terminalState=null, terminatedReason=null}
13:48:33,304 ... DefaultStateTransitionManager [] - Transitioning from Cooldown
to Idling, job a59d5fc22ec6d46289dd2f3064b4e17f.
13:48:33,318 ... Current rescale Rescale{... triggerCause=UPDATE_REQUIREMENT,
terminalState=IGNORED, terminatedReason=NO_RESOURCES_OR_PARALLELISMS_CHANGE} is
null or terminated, so the update action is ignored.
{code}
Nightly 2026-10-08 (test_cron_hadoop343 core, JDK 17),
testRecordRescaleForNewResourcesRequirements:
https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79930&view=logs&j=5af3fd07-293c-5932-0967-2a3127c0f840
{code}
[ERROR]
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.testRecordRescaleForNewResourcesRequirements
-- Time elapsed: 0.494 s <<< FAILURE!
org.opentest4j.AssertionFailedError:
expected: SUCCEEDED
but was: null
at
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCaseBase.assertTerminalRelatedFields(RescaleTimelineITCaseBase.java:151)
at
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.lambda$testRecordRescaleForNewResourcesRequirements$1(RescaleTimelineITCase.java:139)
{code}
Here the job was again first scheduled at parallelism 2 instead of 4, so the
requirement update to 1..2 did not change the parallelism and the
UPDATE_REQUIREMENT rescale was still open when the test took the snapshot. It
ended with JOB_FINISHED after the unblock:
{code}
02:03:07,330 ... Rescale{... rescaleAttemptId=1}, ...
preRescaleParallelism=null, postRescaleParallelism=2 ...
triggerCause=INITIAL_SCHEDULE, terminalState=COMPLETED,
terminatedReason=SUCCEEDED}
02:03:07,417 ... Rescale [] - Updated rescale is: Rescale{...
triggerCause=UPDATE_REQUIREMENT, terminalState=null, terminatedReason=null}
02:03:07,426 ... Rescale{... rescaleAttemptId=1}, ... preRescaleParallelism=2,
postRescaleParallelism=null ... triggerCause=UPDATE_REQUIREMENT,
terminalState=IGNORED, terminatedReason=JOB_FINISHED}
{code}
That is four of the 101 master and release-2.3 runs in the last seven days.
> RescaleTimelineITCase.testRecordRescaleForNewAvailableResource records three
> rescales instead of two
> ----------------------------------------------------------------------------------------------------
>
> Key: FLINK-40789
> URL: https://issues.apache.org/jira/browse/FLINK-40789
> Project: Flink
> Issue Type: Bug
> Components: Runtime / Coordination
> Affects Versions: 2.3.1, 2.4.0
> Reporter: Martijn Visser
> Assignee: Naveen Joseph
> Priority: Critical
> Labels: test-stability
>
> https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79368&view=logs&j=0da23115-68bb-5dcd-192c-bd4c8adebde1
> (master push 2026-09-23, test_ci core, JDK 17)
> {code}
> [ERROR]
> org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.testRecordRescaleForNewAvailableResource
> -- Time elapsed: 0.722 s <<< FAILURE!
> java.lang.AssertionError:
> Expected size: 2 but was: 3 in:
> [Rescale{... rescaleAttemptId=3, ... preRescaleParallelism=4,
> postRescaleParallelism=6 ... triggerCause=NEW_RESOURCE_AVAILABLE,
> terminalState=COMPLETED ...},
> Rescale{... rescaleAttemptId=2, ... preRescaleParallelism=2,
> postRescaleParallelism=4 ... triggerCause=NEW_RESOURCE_AVAILABLE,
> terminalState=COMPLETED ...},
> Rescale{... rescaleAttemptId=1, ... preRescaleParallelism=null,
> postRescaleParallelism=2 ... triggerCause=INITIAL_SCHEDULE,
> terminalState=COMPLETED ...}]
> at
> org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.lambda$testRecordRescaleForNewAvailableResource$2(RescaleTimelineITCase.java:170)
> at
> org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.runAdaptedParameterizedAssertion(RescaleTimelineITCase.java:397)
> at
> org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.testRecordRescaleForNewAvailableResource(RescaleTimelineITCase.java:164)
> {code}
>
>
> The test expects the initial schedule at parallelism 4 and
> one upscale to 6 after the third TaskManager starts. Here the job was first
> scheduled at 2 and went 2 -> 4 -> 6. The other RescaleTimelineITCase methods
> had similar races fixed in FLINK-40010, FLINK-40067 and FLINK-40076; this
> method has no ticket.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)