[
https://issues.apache.org/jira/browse/FLINK-40823?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18119738#comment-18119738
]
sepuri sai krishna commented on FLINK-40823:
--------------------------------------------
I have a reproducer for this and would like to write the doc change. Could I be
assigned ?
> State Processor API docs do not mention that a checkpoint of a running job
> can be deleted while it is being read
> ----------------------------------------------------------------------------------------------------------------
>
> Key: FLINK-40823
> URL: https://issues.apache.org/jira/browse/FLINK-40823
> Project: Flink
> Issue Type: Improvement
> Components: API / State Processor, Documentation
> Affects Versions: 2.3.0
> Reporter: sepuri sai krishna
> Priority: Major
> Attachments: CheckpointReadLifecycleRepro.java, pom.xml
>
>
> The State Processor API page tells users they can read a checkpoint, but
> never mentions that
> a checkpoint belonging to a still-running job can be removed by ordinary
> retention before the
> read finishes.
> Current text:
> {noformat}
> line 29 ... reading, writing, and modifying savepoints and checkpoints ...
> line 82 Reading state begins by specifying the path to a valid savepoint or
> checkpoint
> along with the StateBackend that should be used to restore the data.
> line 414 SavepointReader savepoint = SavepointReader.read(env,
> "hdfs://checkpoint-dir", new HashMapStateBackend());
> {noformat}
> Searching docs/content/docs/libs/state_processor_api.md for "retain",
> "retention",
> "subsum", "num-retained", "delet" and "lifetime" returns zero matches. Same
> for the Chinese
> page.
> What happens in practice, measured on released 2.3.0: a read planned against
> a valid
> checkpoint of a running job fails once that checkpoint is subsumed.
> {noformat}
> producer running, HashMapStateBackend, filesystem storage
> execution.checkpointing.num-retained = 1
> execution.checkpointing.data-inline-threshold = 0 (so state is in real
> files)
> 1. checkpoint C completes, state files present on disk
> 2. SavepointReader.read(...) + readKeyedState(...) -- graph built, job
> not submitted
> 3. producer completes ONE more ordinary checkpoint -> all of C's files are
> gone
> 4. the planned read is executed
> {noformat}
> {noformat}
> java.io.IOException: Failed to restore state backend
> ... Could not restore keyed state backend ...
> ... BackendBuildingException: Failed when trying to restore heap backend
> java.io.FileNotFoundException:
> .../chk-2/9c6fe938-90cb-4d74-be32-1472192ac732 (No such file or directory)
> {noformat}
> The missing file is the one recorded from checkpoint C before the read was
> planned. The
> producer is still RUNNING throughout; nothing is deleted by the test.
> Reproduced 4 out of 4
> runs, each on a fresh cluster and directory, each after exactly one
> subsequent checkpoint.
> Two supported settings avoid it completely, and both were verified in the
> same harness:
> {noformat}
> num-retained = 20 C survived 12 checkpoints,
> read succeeded
> savepoint taken with terminate = false C survived 12 checkpoints,
> read succeeded,
> producer kept running
> {noformat}
> So this is a documentation gap rather than a bug: the behaviour is
> reasonable, but a reader
> following the page has no way to know the window exists or that these two
> mitigations are the
> answer.
> Suggested scope: a short note in the "Reading State" section covering the
> subsumption window
> and the two mitigations, in both the English and Chinese pages.
> A separate question, which I have not filed: the failure surfaces as a bare
> FileNotFoundException nine causes deep, which gives no hint that the cause is
> ordinary
> checkpoint subsumption. Happy to raise that separately if it is worth it.
> Reproducer attached; it runs against released 2.3.0 artifacts and needs no
> Flink checkout.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)