DanielLeens opened a new issue, #12132: URL: https://github.com/apache/seatunnel/issues/12132
### Search before asking - [x] I had searched in the [issues](https://github.com/apache/seatunnel/issues?q=is%3Aissue) and found no similar issues. ### What happened `kudu-connector-it` regularly runs into the 90-minute job limit and is cancelled by GitHub Actions. In every case `KuduIT` is executing a job on a Flink leg, the Flink JobManager stops producing any output right after deploying the Kudu source task, the TaskManager reports the JobManager heartbeat timeout two minutes later, and `container.executeJob` (`flink run` inside the JobManager container) never returns. Sightings (all cancelled at exactly 90 minutes): - apache/seatunnel dev push run 33970546560, `kudu-connector-it (11)`, 14:03 -> 15:34 UTC - PR #11458, fork run `zhangshenghang/seatunnel` 33971790374, `kudu-connector-it (11)`, 15:37 -> 17:07, Flink 1.15.3 leg, test `testKuduTableListWithRegex` - PR #11077, fork run `hesam-oxe/seatunnel` 33971836407, `kudu-connector-it (8)`, 16:10 -> 17:40, Flink 1.18.0 leg, test `testKuduFilter` - PR #11757, fork run `waterWang/seatunnel` 33971820086, `kudu-connector-it (11)`, 16:46 -> 18:17, Flink 1.15.3 leg, test `testKudu` The sibling JDK leg of the same run passed in 26-27 minutes each time. ### Timeline (run 33971820086, job 101321902856) - `16:58:16.552` JobManager (`tyrantlucifer/flink:1.15.3-scala_2.12_hadoop27`): `Deploying Source: Kudu-Source -> MultiTableSink-Sink: Writer (1/1) (attempt #0)` -- this is the last line the JobManager container ever prints. - `16:58:16.594` TaskManager: `Downloading 39a807708be29305ebe397dee12e3406/p-... from jobmanager` (BLOB download for the task JAR), then silence. - `17:00:16` TaskManager: `The heartbeat of JobManager with id ... timed out` (the e2e containers run with `heartbeat.timeout: 120000`), task switched to FAILED. - `17:00:33` TaskManager: `The heartbeat of ResourceManager ... timed out`, then every 10 s `Could not resolve ResourceManager address akka.tcp://flink@jobmanager:6123/user/rpc/resourcemanager_*, retrying in 10000 ms: Could not connect to rpc endpoint`. - `17:05:43` TaskManager: `Terminating TaskManagerRunner with exit code 1` after the registration timeout. - `18:16:59` GitHub cancels the job (`The operation was canceled.`). The test JVM was still blocked in `KuduIT.testKudu -> container.executeJob`. No `OutOfMemoryError`, `hs_err`, or fatal-error banner appears in the captured JobManager/TaskManager output; the JobManager container itself keeps running (otherwise the `docker exec` running `flink run` would have terminated). ### What you expected to happen The Kudu e2e should either complete or fail within minutes. Two things need attention: 1. Why the Flink JobManager becomes unresponsive exactly when the Kudu source task is deployed on the 1.15/1.18 legs (the SeaTunnel source enumerator for Kudu runs inside the JobManager via the operator coordinator). The Kudu containers are started without any memory limit (`apache/kudu:1.15.0` master and tserver default to 80% of host RAM each), which is worth checking against the runner's memory. 2. `AbstractTestFlinkContainer.executeJob` has no upper bound: with the JobManager frozen, the Flink CLI keeps polling and the job burns the whole 90-minute budget. A bounded wait would turn this into a fast, diagnosable failure. ### SeaTunnel Version dev (`af0a647d`) and the three PR heads above; `apache/kudu:1.15.0`, Flink 1.15.3 and 1.18.0 e2e containers. ### Engine Flink ### Additional context Filed while triaging CI for #11757, #11458 and #11077; none of those PRs touch connector-kudu or the Flink e2e containers. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
