DanielLeens commented on issue #12132:
URL: https://github.com/apache/seatunnel/issues/12132#issuecomment-5554347869

   Fifth sighting, with a variant that narrows the failure: PR #11727, fork run 
`abdessalems/seatunnel` 33978895542, job `kudu-connector-it (11)` (job 
101340798483), 18:02 -> 19:32 UTC, cancelled at the 90-minute limit on the 
Flink 1.18.0 leg during `testKuduFilter`.
   
   Here the JobManager container did **not** go silent, so the loop is visible:
   
   - `18:17:32` previous leg finishes; the Flink 1.18 JobManager/TaskManager 
pair starts at 18:17:34.
   - The first job on the fresh pair (`kudu_to_assert.conf`, job `a8b499b1...`) 
is submitted and never reaches RUNNING.
   - `18:19:50` ResourceManager: `The heartbeat of JobManager with id ... timed 
out` (`heartbeat.timeout: 120000`, i.e. exactly two minutes after the job was 
submitted). The JobMaster stopped answering heartbeats while the rest of the 
JVM kept working.
   - From 18:22 until the cancel at 19:32 the same cycle repeats 263 times: 
`FineGrainedSlotManager - Matching resource requirements` / 
`DefaultSlotStatusSyncer - Starting allocation of slot` on the JobManager; `Add 
job ... for job leader monitoring` / `Try to register at job manager 
pekko.tcp://flink@jobmanager:6123/user/rpc/jobmanager_2` / `Could not resolve 
JobManager address` (78x) / `Remove job` / `Free slot` on the TaskManager; plus 
6x `BlobClient - Failed to fetch BLOB`. The job is never failed and `flink run` 
keeps waiting.
   
   Across all five sightings the hang is always the **first job submitted on a 
freshly started Flink JobManager/TaskManager pair** (1.15.3 or 1.18.0), the 
JobMaster stops heartbeating exactly at job start, and the JobMaster-side 
timeouts fire 120 s later. In Flink, `SourceCoordinator#start` calls 
`source.createEnumerator(context)` synchronously on the JobMaster main thread, 
so a Kudu client call that blocks or retries without bound inside 
`KuduSource#createEnumerator` / the split enumerator constructor would produce 
exactly this picture: the JobMaster main thread is stuck, its heartbeats stop, 
the slot request loop keeps spinning from the ResourceManager side, and the job 
can neither run nor fail. That is the first place to look (a JobManager thread 
dump from a reproducing run would confirm it).


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to