DanielLeens commented on issue #12132: URL: https://github.com/apache/seatunnel/issues/12132#issuecomment-5554347869
Fifth sighting, with a variant that narrows the failure: PR #11727, fork run `abdessalems/seatunnel` 33978895542, job `kudu-connector-it (11)` (job 101340798483), 18:02 -> 19:32 UTC, cancelled at the 90-minute limit on the Flink 1.18.0 leg during `testKuduFilter`. Here the JobManager container did **not** go silent, so the loop is visible: - `18:17:32` previous leg finishes; the Flink 1.18 JobManager/TaskManager pair starts at 18:17:34. - The first job on the fresh pair (`kudu_to_assert.conf`, job `a8b499b1...`) is submitted and never reaches RUNNING. - `18:19:50` ResourceManager: `The heartbeat of JobManager with id ... timed out` (`heartbeat.timeout: 120000`, i.e. exactly two minutes after the job was submitted). The JobMaster stopped answering heartbeats while the rest of the JVM kept working. - From 18:22 until the cancel at 19:32 the same cycle repeats 263 times: `FineGrainedSlotManager - Matching resource requirements` / `DefaultSlotStatusSyncer - Starting allocation of slot` on the JobManager; `Add job ... for job leader monitoring` / `Try to register at job manager pekko.tcp://flink@jobmanager:6123/user/rpc/jobmanager_2` / `Could not resolve JobManager address` (78x) / `Remove job` / `Free slot` on the TaskManager; plus 6x `BlobClient - Failed to fetch BLOB`. The job is never failed and `flink run` keeps waiting. Across all five sightings the hang is always the **first job submitted on a freshly started Flink JobManager/TaskManager pair** (1.15.3 or 1.18.0), the JobMaster stops heartbeating exactly at job start, and the JobMaster-side timeouts fire 120 s later. In Flink, `SourceCoordinator#start` calls `source.createEnumerator(context)` synchronously on the JobMaster main thread, so a Kudu client call that blocks or retries without bound inside `KuduSource#createEnumerator` / the split enumerator constructor would produce exactly this picture: the JobMaster main thread is stuck, its heartbeats stop, the slot request loop keeps spinning from the ResourceManager side, and the job can neither run nor fail. That is the first place to look (a JobManager thread dump from a reproducing run would confirm it). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
