DanielLeens commented on issue #12132: URL: https://github.com/apache/seatunnel/issues/12132#issuecomment-5605332907
Sightings update from job logs, 2026-09-06 to 2026-09-09: - #12130 (`awsomesud347/seatunnel` run 34089819303): `kudu-connector-it (11)` cancelled at 90 minutes in all four attempts. Attempt 1 on Flink 1.15.3 (`testKuduMultipleReadWithRegex`: TaskManager heartbeat timeout 06:48:06, TaskManager exit 06:53:21, then silence), attempt 4 on Flink 1.18.0 (`testKuduMultipleRead`: JobMaster `Deploying ... (attempt #0)` at 19:39:48 is its last line, TaskManager `Slot offering to JobManager did not finish in time` at 19:39:58, ResourceManager heartbeat timeout at 19:41:48, slot-allocation loop until the cancel). The previous head of the same PR lost the JDK 8 leg the same way (run 34001671707 attempt 1, Flink 1.18.0, `testKuduWholeDatabaseRead`, `Slot offering ... did not finish in time` 10 s after deploy). - #12203 (`DanielLeens/seatunnel` run 34207958769): both legs, Flink 1.18.0 (`testKudu` on JDK 11, `testKuduMultipleRead` on JDK 8). - #12169 (`JeremyXin/seatunnel` run 34333231463): JDK 8 leg, Flink 1.15.3. #12165 (`Vivek1106-04/seatunnel` run 34111028888): JDK 8 leg, Flink 1.18.0. - apache scheduled dev run 34044345096 (2026-09-06): both legs, on the Flink **1.16.0** (JDK 11, job 101516951936) and **1.17.2** (JDK 8, job 101516952135) legs of the full matrix. So the freeze is not limited to 1.15.3 and 1.18.0; it has still not been seen on 1.13.6 or 1.20.1. Across the PR runs of 2026-09-06 to 2026-09-09 that executed the Kudu job I count 24 green Kudu jobs against 7 cancelled ones, so roughly one Kudu leg in four hangs, on JDK 8 and JDK 11 alike. One detail that is consistent in every log where the TaskManager side is visible: `Slot offering to JobManager did not finish in time. Retrying the slot offering.` appears 10 s after `Deploying ... (attempt #0)`, and the `BlobClient - Downloading <job>/p-... from jobmanager` started at the same second never completes. So the JobMaster RPC endpoint is already unresponsive within 10 s of `Execution.deploy`, at the same instant the BlobServer stops serving. The JobManager thread dump remains the missing piece; a bounded wait in `AbstractTestFlinkContainer.executeJob` that dumps the JobManager threads (`jstack` via `docker exec`) before failing would capture it on the next occurrence instead of burning 90 minutes. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
