[ 
https://issues.apache.org/jira/browse/FLINK-40663?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121536#comment-18121536
 ] 

Martijn Visser commented on FLINK-40663:
----------------------------------------

One more on release-1.20, nightly 2026-10-01 (Java 11, module table):
https://github.com/apache/flink/actions/runs/36804869542/job/110195862980

{code}
[ERROR] 
org.apache.flink.table.planner.runtime.stream.sql.WindowDistinctAggregateStressITCase.testCumulateWindowRollupUnderBackpressure
 -- Time elapsed: 624.5 s <<< ERROR!
java.util.concurrent.ExecutionException: 
org.apache.flink.table.api.TableException: Failed to wait job finish
        at 
org.apache.flink.table.planner.runtime.stream.sql.WindowDistinctAggregateStressITCase.runQuery(WindowDistinctAggregateStressITCase.java:191)
        at 
org.apache.flink.table.planner.runtime.stream.sql.WindowDistinctAggregateStressITCase.testCumulateWindowRollupUnderBackpressure(WindowDistinctAggregateStressITCase.java:149)
...
Caused by: org.apache.flink.runtime.JobException: Recovery is suppressed by 
FixedDelayRestartBackoffTimeStrategy(maxNumberRestartAttempts=1, 
backoffTimeMS=0)
...
Caused by: org.apache.flink.util.FlinkRuntimeException: Exceeded checkpoint 
tolerable failure threshold. The latest checkpoint failed due to Checkpoint 
expired before completing., view the Checkpoint History tab or the Job Manager 
log to find out why continuous checkpoints failed.
{code}

The fork log shows why the checkpoint never completes. After the test's 
artificial failure and restore, the JobManager fails to process an 
acknowledgement of checkpoint 11. The checkpoint stays pending until it expires 
ten minutes later, and the job then fails because its one restart is already 
used:

{code}
03:20:39,724 [    Checkpoint Timer] INFO  
org.apache.flink.runtime.checkpoint.CheckpointCoordinator    [] - Triggering 
checkpoint 11 (type=CheckpointType{name='Checkpoint', 
sharingFilesStrategy=FORWARD_BACKWARD}) @ 1790824839723 for job 
f491f526d2be204071333cb1c659ebad.
03:20:39,784 [jobmanager-io-thread-4] WARN  
org.apache.flink.runtime.jobmaster.JobMaster                 [] - Error while 
processing AcknowledgeCheckpoint message
java.lang.IllegalStateException: Attempt to reference unknown state: 
f71696db-7f3f-37f7-8d7e-0edbb2a2ba7c
        at 
org.apache.flink.util.Preconditions.checkState(Preconditions.java:193)
        at 
org.apache.flink.runtime.state.SharedStateRegistryImpl.registerReference(SharedStateRegistryImpl.java:97)
        at 
org.apache.flink.runtime.state.SharedStateRegistry.registerReference(SharedStateRegistry.java:53)
        at 
org.apache.flink.runtime.state.IncrementalRemoteKeyedStateHandle.registerSharedStates(IncrementalRemoteKeyedStateHandle.java:289)
        ...
        at 
org.apache.flink.runtime.checkpoint.CheckpointCoordinator.receiveAcknowledgeMessage(CheckpointCoordinator.java:1245)
03:30:39,724 [    Checkpoint Timer] INFO  
org.apache.flink.runtime.checkpoint.CheckpointCoordinator    [] - Checkpoint 11 
of job f491f526d2be204071333cb1c659ebad expired before completing.
{code}

The fork logs of the five earlier failed jobs since 2026-09-24 show the same 
exception, each followed by the same expiry. FLINK-38574 describes this 
exception for RocksDB incremental checkpoints. Its fix is on master, 
release-2.3 and release-2.2 and was backported to 1.19, 2.0 and 2.1, but not to 
release-1.20, the only branch where this test fails. Three of the seven GitHub 
Actions nightlies on release-1.20 in the last week hit it, and none of the runs 
on the other branches.

> WindowDistinctAggregateStressITCase.testCumulateWindowRollupUnderBackpressure 
> times out after 600s 
> ---------------------------------------------------------------------------------------------------
>
>                 Key: FLINK-40663
>                 URL: https://issues.apache.org/jira/browse/FLINK-40663
>             Project: Flink
>          Issue Type: Bug
>          Components: Table SQL / Runtime
>    Affects Versions: 1.20.6
>            Reporter: Martijn Visser
>            Priority: Critical
>              Labels: test-stability
>
> {code}
> [ERROR] 
> org.apache.flink.table.planner.runtime.stream.sql.WindowDistinctAggregateStressITCase.testCumulateWindowRollupUnderBackpressure
>  -- Time elapsed: 617.2 s <<< ERROR!
> java.util.concurrent.ExecutionException: 
> org.apache.flink.table.api.TableException: Failed to wait job finish
>         at 
> java.base/java.util.concurrent.CompletableFuture.reportGet(CompletableFuture.java:395)
>         at 
> org.apache.flink.table.api.internal.TableResultImpl.awaitInternal(TableResultImpl.java:122)
>         at 
> org.apache.flink.table.planner.runtime.stream.sql.WindowDistinctAggregateStressITCase.runQuery(WindowDistinctAggregateStressITCase.java:191)
>         at 
> org.apache.flink.table.planner.runtime.stream.sql.WindowDistinctAggregateStressITCase.testCumulateWindowRollupUnderBackpressure(WindowDistinctAggregateStressITCase.java:149)
> Caused by: org.apache.flink.table.api.TableException: Failed to wait job 
> finish
>         at 
> org.apache.flink.table.api.internal.InsertResultProvider.hasNext(InsertResultProvider.java:85)
> {code}
> Three occurrences in the window, on two different profiles and on both a 
> nightly and a push run:
> https://github.com/apache/flink/actions/runs/34919682822 (nightly 2026-09-15, 
> Java 11)
> https://github.com/apache/flink/actions/runs/34798014224 (nightly 2026-09-14, 
> Java 17)
> https://github.com/apache/flink/actions/runs/34861606726 (push 2026-09-14, 
> Java 8)



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to