isaac-montenegro-jimenez opened a new pull request, #29409:
URL: https://github.com/apache/flink/pull/29409
## What is the purpose of the change
`CheckpointCoordinator` creates the checkpoint base directories (`shared/`,
`taskowned/`) in its constructor when periodic checkpointing is configured. On
file systems where `mkdirs` is a real remote call (e.g. GCS, ADLS Gen2), this
blocks ExecutionGraph creation and adds roughly 0.1–0.5 s to every job start.
This PR adds `execution.checkpointing.initialize-base-locations.eagerly`
(default `true`, i.e. unchanged behavior). When set to `false`, the base
locations are created on the IO executor when the first checkpoint is triggered.
It also fixes the existing lazy-initialization path, which marked base
locations as initialized before the initialization ran.
## Brief change log
- [runtime] Fix premature base-location flag in `CheckpointCoordinator`:
`baseLocationsForCheckpointInitialized` was set when a trigger started. If the
first trigger was a savepoint, or the first initialization failed, base
locations were never initialized for later checkpoints. The flag is now set
only after `initializeBaseLocationsForCheckpoint()` succeeds, and is `volatile`
since it is written on the IO executor.
- [runtime] Add `CheckpointingOptions.INITIALIZE_BASE_LOCATIONS_EAGERLY`,
forwarded through `StreamGraph` and `CheckpointCoordinatorConfiguration` to the
coordinator. The lazy path is logged at INFO.
- Regenerated the config docs.
## Verifying this change
This change added tests and can be verified as follows:
- `CheckpointCoordinatorTest`: first checkpoint after a savepoint
initializes base locations; initialization is retried after a failure and not
repeated after success; eager initialization is not repeated on trigger; with
the option disabled, initialization happens on the first checkpoint only. The
first two fail without the fix.
- `StreamGraphGeneratorTest`: the option reaches
`CheckpointCoordinatorConfiguration`.
- Ran the full `flink-core`, `flink-runtime`, `flink-streaming-java` and
`flink-docs` builds (`./mvnw verify`), and `CheckpointFailureManagerITCase`,
`UnalignedCheckpointFailureHandlingITCase`, `StateBackendITCase` in
`flink-tests`.
## Does this pull request potentially affect one of the following parts:
- Dependencies (does it add or upgrade a dependency): no
- The public API, i.e., is any changed class annotated with
`@Public(Evolving)`: yes (a new option in the `@PublicEvolving`
`CheckpointingOptions`)
- The serializers: no
- The runtime per-record code paths (performance sensitive): no
- Anything that affects deployment or recovery: JobManager (and its
components), Checkpointing, Kubernetes/Yarn, ZooKeeper: yes (checkpointing; see
notes below)
- The S3 file system connector: no
## Documentation
- Does this pull request introduce a new feature? yes
- If yes, how is the feature documented? docs (generated config reference)
## Notes for reviewers
- The default is unchanged. With `false`, an unusable checkpoint location is
no longer detected when the job starts; it surfaces as checkpoint failures and
fails the job once the tolerable failure count is exceeded.
- Upstream deliberately kept eager initialization for periodic checkpointing
(apache/flink#17278, "let's do it twofold") to fail early. The option keeps
that as the default.
- The first commit changes behavior even with the option untouched, in two
cases that were broken before: a savepoint as the first trigger, and a failed
first initialization.
---
##### Was generative AI tooling used to co-author this PR?
- [X] Yes (please specify the tool below)
Generated-by: Claude Code (claude-sonnet-5-5)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]