[
https://issues.apache.org/jira/browse/FLINK-40663?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121536#comment-18121536
]
Martijn Visser commented on FLINK-40663:
----------------------------------------
One more on release-1.20, nightly 2026-10-01 (Java 11, module table):
https://github.com/apache/flink/actions/runs/36804869542/job/110195862980
{code}
[ERROR]
org.apache.flink.table.planner.runtime.stream.sql.WindowDistinctAggregateStressITCase.testCumulateWindowRollupUnderBackpressure
-- Time elapsed: 624.5 s <<< ERROR!
java.util.concurrent.ExecutionException:
org.apache.flink.table.api.TableException: Failed to wait job finish
at
org.apache.flink.table.planner.runtime.stream.sql.WindowDistinctAggregateStressITCase.runQuery(WindowDistinctAggregateStressITCase.java:191)
at
org.apache.flink.table.planner.runtime.stream.sql.WindowDistinctAggregateStressITCase.testCumulateWindowRollupUnderBackpressure(WindowDistinctAggregateStressITCase.java:149)
...
Caused by: org.apache.flink.runtime.JobException: Recovery is suppressed by
FixedDelayRestartBackoffTimeStrategy(maxNumberRestartAttempts=1,
backoffTimeMS=0)
...
Caused by: org.apache.flink.util.FlinkRuntimeException: Exceeded checkpoint
tolerable failure threshold. The latest checkpoint failed due to Checkpoint
expired before completing., view the Checkpoint History tab or the Job Manager
log to find out why continuous checkpoints failed.
{code}
The fork log shows why the checkpoint never completes. After the test's
artificial failure and restore, the JobManager fails to process an
acknowledgement of checkpoint 11. The checkpoint stays pending until it expires
ten minutes later, and the job then fails because its one restart is already
used:
{code}
03:20:39,724 [ Checkpoint Timer] INFO
org.apache.flink.runtime.checkpoint.CheckpointCoordinator [] - Triggering
checkpoint 11 (type=CheckpointType{name='Checkpoint',
sharingFilesStrategy=FORWARD_BACKWARD}) @ 1790824839723 for job
f491f526d2be204071333cb1c659ebad.
03:20:39,784 [jobmanager-io-thread-4] WARN
org.apache.flink.runtime.jobmaster.JobMaster [] - Error while
processing AcknowledgeCheckpoint message
java.lang.IllegalStateException: Attempt to reference unknown state:
f71696db-7f3f-37f7-8d7e-0edbb2a2ba7c
at
org.apache.flink.util.Preconditions.checkState(Preconditions.java:193)
at
org.apache.flink.runtime.state.SharedStateRegistryImpl.registerReference(SharedStateRegistryImpl.java:97)
at
org.apache.flink.runtime.state.SharedStateRegistry.registerReference(SharedStateRegistry.java:53)
at
org.apache.flink.runtime.state.IncrementalRemoteKeyedStateHandle.registerSharedStates(IncrementalRemoteKeyedStateHandle.java:289)
...
at
org.apache.flink.runtime.checkpoint.CheckpointCoordinator.receiveAcknowledgeMessage(CheckpointCoordinator.java:1245)
03:30:39,724 [ Checkpoint Timer] INFO
org.apache.flink.runtime.checkpoint.CheckpointCoordinator [] - Checkpoint 11
of job f491f526d2be204071333cb1c659ebad expired before completing.
{code}
The fork logs of the five earlier failed jobs since 2026-09-24 show the same
exception, each followed by the same expiry. FLINK-38574 describes this
exception for RocksDB incremental checkpoints. Its fix is on master,
release-2.3 and release-2.2 and was backported to 1.19, 2.0 and 2.1, but not to
release-1.20, the only branch where this test fails. Three of the seven GitHub
Actions nightlies on release-1.20 in the last week hit it, and none of the runs
on the other branches.
> WindowDistinctAggregateStressITCase.testCumulateWindowRollupUnderBackpressure
> times out after 600s
> ---------------------------------------------------------------------------------------------------
>
> Key: FLINK-40663
> URL: https://issues.apache.org/jira/browse/FLINK-40663
> Project: Flink
> Issue Type: Bug
> Components: Table SQL / Runtime
> Affects Versions: 1.20.6
> Reporter: Martijn Visser
> Priority: Critical
> Labels: test-stability
>
> {code}
> [ERROR]
> org.apache.flink.table.planner.runtime.stream.sql.WindowDistinctAggregateStressITCase.testCumulateWindowRollupUnderBackpressure
> -- Time elapsed: 617.2 s <<< ERROR!
> java.util.concurrent.ExecutionException:
> org.apache.flink.table.api.TableException: Failed to wait job finish
> at
> java.base/java.util.concurrent.CompletableFuture.reportGet(CompletableFuture.java:395)
> at
> org.apache.flink.table.api.internal.TableResultImpl.awaitInternal(TableResultImpl.java:122)
> at
> org.apache.flink.table.planner.runtime.stream.sql.WindowDistinctAggregateStressITCase.runQuery(WindowDistinctAggregateStressITCase.java:191)
> at
> org.apache.flink.table.planner.runtime.stream.sql.WindowDistinctAggregateStressITCase.testCumulateWindowRollupUnderBackpressure(WindowDistinctAggregateStressITCase.java:149)
> Caused by: org.apache.flink.table.api.TableException: Failed to wait job
> finish
> at
> org.apache.flink.table.api.internal.InsertResultProvider.hasNext(InsertResultProvider.java:85)
> {code}
> Three occurrences in the window, on two different profiles and on both a
> nightly and a push run:
> https://github.com/apache/flink/actions/runs/34919682822 (nightly 2026-09-15,
> Java 11)
> https://github.com/apache/flink/actions/runs/34798014224 (nightly 2026-09-14,
> Java 17)
> https://github.com/apache/flink/actions/runs/34861606726 (push 2026-09-14,
> Java 8)
--
This message was sent by Atlassian Jira
(v8.20.10#820010)