MartijnVisser opened a new pull request, #29322:
URL: https://github.com/apache/flink/pull/29322

   ## What is the purpose of the change
   
   `LocalRecoveryTest.testStateSizeIsConsideredForLocalRecoveryOnRestart` also 
fails with an NPE on `targetLocation` in 
`PendingCheckpoint.finalizeCheckpoint`, a path the two earlier fixes for this 
ticket do not touch. Seen twice on release-2.2 so far, master has the same code:
   
   
https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79477&view=logs&j=d89de3df-4600-5585-dadc-9bbc9a5e661c
 (nightly 2026-09-27, `test_cron_hadoop313 core`)
   
https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=78771&view=logs&j=d06b80b4-9e88-5d40-12a2-18072cf60528
 (nightly 2026-09-06, `test_cron_jdk21 core`)
   
   The test acknowledges the checkpoint once it counts as in progress, which is 
before the coordinator sets its storage location.
   
   ## Brief change log
   
     - `LocalRecoveryTest` waits until the checkpoint was triggered on all four 
tasks before acknowledging it, like 
`DefaultSchedulerTest#getCheckpointTriggeredLatch`. Tasks are only triggered 
after the storage location is set.
     - `AdaptiveSchedulerTestBase#getSlotPoolWithFreeSlots` gets an overload 
that takes the slots' `TaskManagerGateway`.
     - `SchedulerTestingUtils#waitForCheckpointInProgress` has no other caller 
and is removed.
   
   ## Verifying this change
   
   This change is already covered by existing tests, such as 
`LocalRecoveryTest`. Measured locally on JDK 17 with assertions enabled, as in 
surefire:
   
     - Holding the checkpoint timer thread for 2 s right after it registers the 
checkpoint fails the test with the CI stack in 100 of 100 runs on master 
(5af0f51cfe4). With this change 0 of 100 fail, and with the old wait put back 
100 of 100 fail again.
     - Without the hold, in a Linux container on one CPU, 44 of 6,000 runs fail 
on master: 43 with this NPE, and one on the 10 minute checkpoint timeout, 
likely because the acknowledgements came before the coordinator added the 
checkpoint to `pendingCheckpoints`. With this change 0 of 6,000. On macOS under 
CPU load it is 34 of 60,000 against 0 of 60,000.
     - The `scheduler.adaptive` tests, `spotless:check` and `checkstyle:check` 
pass.
   
   ## Does this pull request potentially affect one of the following parts:
   
     - Dependencies (does it add or upgrade a dependency): no
     - The public API, i.e., is any changed class annotated with 
`@Public(Evolving)`: no
     - The serializers: no
     - The runtime per-record code paths (performance sensitive): no
     - Anything that affects deployment or recovery: JobManager (and its 
components), Checkpointing, Kubernetes/Yarn, ZooKeeper: no
     - The S3 file system connector: no
   
   ## Documentation
   
     - Does this pull request introduce a new feature? no
     - If yes, how is the feature documented? not applicable
   
   ---
   
   ##### Was generative AI tooling used to co-author this PR?
   
   - [X] Yes (please specify the tool below)
   
   Generated-by: Claude Code (Claude Opus 5.5)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to