[
https://issues.apache.org/jira/browse/FLINK-40789?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18123272#comment-18123272
]
Martijn Visser commented on FLINK-40789:
----------------------------------------
Two more, the release-2.3 and master nightlies of 2026-10-04 (AdaptiveScheduler
profile, JDK 17):
https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79753&view=logs&j=0e7be18f-84f2-53f0-a32d-4a5e4a174679
(release-2.3)
https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79750&view=logs&j=0e7be18f-84f2-53f0-a32d-4a5e4a174679
(master, RescaleTimelineHistoryEnabledITCase)
{code}
[ERROR]
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.testRescaleTerminatedByJobCancelled
-- Time elapsed: 10.63 s <<< ERROR!
java.util.concurrent.TimeoutException: Condition was not met within 10000 ms.
at
org.apache.flink.core.testutils.CommonTestUtils.waitUtil(CommonTestUtils.java:218)
at
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.waitUntilConditionWithTimeout(RescaleTimelineITCase.java:698)
at
org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.testRescaleTerminatedByJobCancelled(RescaleTimelineITCase.java:325)
{code}
On release-2.3 the job was again first scheduled at parallelism 2, so there are
three rescales:
{code}
Rescale{... rescaleAttemptId=1}, ... postRescaleParallelism=2 ...
triggerCause=INITIAL_SCHEDULE, terminalState=COMPLETED,
terminatedReason=SUCCEEDED}
Rescale{... rescaleAttemptId=2}, ... preRescaleParallelism=2 ...
triggerCause=NEW_RESOURCE_AVAILABLE, terminalState=IGNORED,
terminatedReason=RESOURCE_REQUIREMENTS_UPDATED}
Rescale{... rescaleAttemptId=1}, ... preRescaleParallelism=2 ...
triggerCause=UPDATE_REQUIREMENT, terminalState=IGNORED,
terminatedReason=JOB_CANCELED}
{code}
Master is the 2026-09-07 shape from my last comment: two rescales, the second
ended with NO_RESOURCES_OR_PARALLELISMS_CHANGE when the cooldown ended, right
before the cancel:
{code}
01:06:10,140 ... DefaultStateTransitionManager [] - Transitioning from Cooldown
to Idling, job 5db28c624a59470afec441d92439a75d.
01:06:10,140 ... Current rescale Rescale{... rescaleAttemptId=1}, ...
preRescaleParallelism=4 ... triggerCause=UPDATE_REQUIREMENT,
terminalState=IGNORED, terminatedReason=NO_RESOURCES_OR_PARALLELISMS_CHANGE} is
null or terminated, so the update action is ignored.
01:06:10,140 ... ExecutionGraph [] - Job Unnamed job
(5db28c624a59470afec441d92439a75d) switched from state RUNNING to CANCELLING.
{code}
That is three of the 104 master and release-2.3 runs in the last seven days,
all in this method.
> RescaleTimelineITCase.testRecordRescaleForNewAvailableResource records three
> rescales instead of two
> ----------------------------------------------------------------------------------------------------
>
> Key: FLINK-40789
> URL: https://issues.apache.org/jira/browse/FLINK-40789
> Project: Flink
> Issue Type: Bug
> Components: Runtime / Coordination
> Affects Versions: 2.3.1, 2.4.0
> Reporter: Martijn Visser
> Priority: Major
> Labels: test-stability
>
> https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79368&view=logs&j=0da23115-68bb-5dcd-192c-bd4c8adebde1
> (master push 2026-09-23, test_ci core, JDK 17)
> {code}
> [ERROR]
> org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.testRecordRescaleForNewAvailableResource
> -- Time elapsed: 0.722 s <<< FAILURE!
> java.lang.AssertionError:
> Expected size: 2 but was: 3 in:
> [Rescale{... rescaleAttemptId=3, ... preRescaleParallelism=4,
> postRescaleParallelism=6 ... triggerCause=NEW_RESOURCE_AVAILABLE,
> terminalState=COMPLETED ...},
> Rescale{... rescaleAttemptId=2, ... preRescaleParallelism=2,
> postRescaleParallelism=4 ... triggerCause=NEW_RESOURCE_AVAILABLE,
> terminalState=COMPLETED ...},
> Rescale{... rescaleAttemptId=1, ... preRescaleParallelism=null,
> postRescaleParallelism=2 ... triggerCause=INITIAL_SCHEDULE,
> terminalState=COMPLETED ...}]
> at
> org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.lambda$testRecordRescaleForNewAvailableResource$2(RescaleTimelineITCase.java:170)
> at
> org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.runAdaptedParameterizedAssertion(RescaleTimelineITCase.java:397)
> at
> org.apache.flink.runtime.scheduler.adaptive.timeline.RescaleTimelineITCase.testRecordRescaleForNewAvailableResource(RescaleTimelineITCase.java:164)
> {code}
>
>
> The test expects the initial schedule at parallelism 4 and
> one upscale to 6 after the third TaskManager starts. Here the job was first
> scheduled at 2 and went 2 -> 4 -> 6. The other RescaleTimelineITCase methods
> had similar races fixed in FLINK-40010, FLINK-40067 and FLINK-40076; this
> method has no ticket.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)