DanielLeens opened a new issue, #12132:
URL: https://github.com/apache/seatunnel/issues/12132

   ### Search before asking
   
   - [x] I had searched in the 
[issues](https://github.com/apache/seatunnel/issues?q=is%3Aissue) and found no 
similar issues.
   
   ### What happened
   
   `kudu-connector-it` regularly runs into the 90-minute job limit and is 
cancelled by GitHub Actions. In every case `KuduIT` is executing a job on a 
Flink leg, the Flink JobManager stops producing any output right after 
deploying the Kudu source task, the TaskManager reports the JobManager 
heartbeat timeout two minutes later, and `container.executeJob` (`flink run` 
inside the JobManager container) never returns.
   
   Sightings (all cancelled at exactly 90 minutes):
   
   - apache/seatunnel dev push run 33970546560, `kudu-connector-it (11)`, 14:03 
-> 15:34 UTC
   - PR #11458, fork run `zhangshenghang/seatunnel` 33971790374, 
`kudu-connector-it (11)`, 15:37 -> 17:07, Flink 1.15.3 leg, test 
`testKuduTableListWithRegex`
   - PR #11077, fork run `hesam-oxe/seatunnel` 33971836407, `kudu-connector-it 
(8)`, 16:10 -> 17:40, Flink 1.18.0 leg, test `testKuduFilter`
   - PR #11757, fork run `waterWang/seatunnel` 33971820086, `kudu-connector-it 
(11)`, 16:46 -> 18:17, Flink 1.15.3 leg, test `testKudu`
   
   The sibling JDK leg of the same run passed in 26-27 minutes each time.
   
   ### Timeline (run 33971820086, job 101321902856)
   
   - `16:58:16.552` JobManager 
(`tyrantlucifer/flink:1.15.3-scala_2.12_hadoop27`): `Deploying Source: 
Kudu-Source -> MultiTableSink-Sink: Writer (1/1) (attempt #0)` -- this is the 
last line the JobManager container ever prints.
   - `16:58:16.594` TaskManager: `Downloading 
39a807708be29305ebe397dee12e3406/p-... from jobmanager` (BLOB download for the 
task JAR), then silence.
   - `17:00:16` TaskManager: `The heartbeat of JobManager with id ... timed 
out` (the e2e containers run with `heartbeat.timeout: 120000`), task switched 
to FAILED.
   - `17:00:33` TaskManager: `The heartbeat of ResourceManager ... timed out`, 
then every 10 s `Could not resolve ResourceManager address 
akka.tcp://flink@jobmanager:6123/user/rpc/resourcemanager_*, retrying in 10000 
ms: Could not connect to rpc endpoint`.
   - `17:05:43` TaskManager: `Terminating TaskManagerRunner with exit code 1` 
after the registration timeout.
   - `18:16:59` GitHub cancels the job (`The operation was canceled.`). The 
test JVM was still blocked in `KuduIT.testKudu -> container.executeJob`.
   
   No `OutOfMemoryError`, `hs_err`, or fatal-error banner appears in the 
captured JobManager/TaskManager output; the JobManager container itself keeps 
running (otherwise the `docker exec` running `flink run` would have terminated).
   
   ### What you expected to happen
   
   The Kudu e2e should either complete or fail within minutes. Two things need 
attention:
   
   1. Why the Flink JobManager becomes unresponsive exactly when the Kudu 
source task is deployed on the 1.15/1.18 legs (the SeaTunnel source enumerator 
for Kudu runs inside the JobManager via the operator coordinator). The Kudu 
containers are started without any memory limit (`apache/kudu:1.15.0` master 
and tserver default to 80% of host RAM each), which is worth checking against 
the runner's memory.
   2. `AbstractTestFlinkContainer.executeJob` has no upper bound: with the 
JobManager frozen, the Flink CLI keeps polling and the job burns the whole 
90-minute budget. A bounded wait would turn this into a fast, diagnosable 
failure.
   
   ### SeaTunnel Version
   
   dev (`af0a647d`) and the three PR heads above; `apache/kudu:1.15.0`, Flink 
1.15.3 and 1.18.0 e2e containers.
   
   ### Engine
   
   Flink
   
   ### Additional context
   
   Filed while triaging CI for #11757, #11458 and #11077; none of those PRs 
touch connector-kudu or the Flink e2e containers.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to