DanielLeens commented on PR #11814: URL: https://github.com/apache/seatunnel/pull/11814#issuecomment-5379439601
Thanks for chasing this down carefully instead of just retrying blindly - that's the right instinct given it's the one test file this PR touches. I pulled the raw logs for all four fork attempts on this head (run 31960889724, attempts 1-4) myself to check your read. Confirming what you found: in every attempt, `FakeSourceToConsoleWithEventReportIT` fails at `SeaTunnelContainer.createSeaTunnelServer` with `ContainerLaunchException: Timed out waiting for log output matching '.*received new worker register:.*'` after ~63-64s, and the captured container log stream has **zero lines** in that window - not even a JVM startup banner, let alone a Hazelcast/worker-register line. Two things stood out when I compared this against the rest of the same job: 1. Elsewhere in the exact same `engine-v2-it` job, the identical `seatunnelhub/openjdk:8u342` image starts and logs "received new worker register" successfully **147 times**, each taking about 6-8 seconds. So the image itself, the Docker pull, and the base container-boot path are all clearly working fine on this runner during this run - it's isolated to this one test. 2. I read `createHttpClient()` in your current `JobEventHttpReportHandler` (lines ~251-263): it's a pure in-memory `OkHttpClient.Builder().build()` call - no DNS, no proxy probing, no socket I/O at construction time. So I don't think your new code is *itself* capable of blocking the container's JVM for 60+ seconds. If the extra okhttp3/kotlin-stdlib classes on the classpath were slowing things down, I'd expect to still see early JVM/log4j startup output before any hang, not total silence for the entire window. That combination (zero output, not partial-then-stall) reads to me as more consistent with the container/runner not getting scheduled promptly at the Docker-daemon level (resource contention on the runner, similar in spirit to the Windows Hazelcast "Node failed to start!" flake and the Couchbase container-bootstrap timeouts we've seen intermittently on unrelated PRs) than with a code-level regression in this PR's diff - but I want to be upfront that I can't fully rule your diff out from static/log analysis alone, since I don't have a `dev`-branch control run of this exact test on this exact runner to compare against. Concrete next step I'd suggest, cheapest first: temporarily bump the wait-strategy timeout for just this one IT (or attach a log consumer from container start rather than only from the wait strategy) and do one more fork run. If the container eventually does emit a startup banner and register message, just later than 63s, that confirms plain scheduling/resource contention and a rerun or a slightly longer timeout is all that's needed - no diff change required. If it stays completely silent even with a much longer window, that would point at something more structural (networking/image) and would be worth flagging separately from this PR's dependency change, since your code path doesn't do anything at construction time that should cause total silence. This is not a known, previously-documented flake signature for this specific test on my side, so I can't just tell you "yes, that's a known issue, ignore it" - but based on the above I also don't think it's your fix that's causing it. I'd treat one more attempt with better diagnostics as the fastest way to get a real answer rather than guessing further from logs alone. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
