DanielLeens commented on PR #11814:
URL: https://github.com/apache/seatunnel/pull/11814#issuecomment-5379439601

   Thanks for chasing this down carefully instead of just retrying blindly - 
that's the right instinct given it's the one test file this PR touches.
   
   I pulled the raw logs for all four fork attempts on this head (run 
31960889724, attempts 1-4) myself to check your read. Confirming what you 
found: in every attempt, `FakeSourceToConsoleWithEventReportIT` fails at 
`SeaTunnelContainer.createSeaTunnelServer` with `ContainerLaunchException: 
Timed out waiting for log output matching '.*received new worker register:.*'` 
after ~63-64s, and the captured container log stream has **zero lines** in that 
window - not even a JVM startup banner, let alone a Hazelcast/worker-register 
line.
   
   Two things stood out when I compared this against the rest of the same job:
   
   1. Elsewhere in the exact same `engine-v2-it` job, the identical 
`seatunnelhub/openjdk:8u342` image starts and logs "received new worker 
register" successfully **147 times**, each taking about 6-8 seconds. So the 
image itself, the Docker pull, and the base container-boot path are all clearly 
working fine on this runner during this run - it's isolated to this one test.
   
   2. I read `createHttpClient()` in your current `JobEventHttpReportHandler` 
(lines ~251-263): it's a pure in-memory `OkHttpClient.Builder().build()` call - 
no DNS, no proxy probing, no socket I/O at construction time. So I don't think 
your new code is *itself* capable of blocking the container's JVM for 60+ 
seconds. If the extra okhttp3/kotlin-stdlib classes on the classpath were 
slowing things down, I'd expect to still see early JVM/log4j startup output 
before any hang, not total silence for the entire window.
   
   That combination (zero output, not partial-then-stall) reads to me as more 
consistent with the container/runner not getting scheduled promptly at the 
Docker-daemon level (resource contention on the runner, similar in spirit to 
the Windows Hazelcast "Node failed to start!" flake and the Couchbase 
container-bootstrap timeouts we've seen intermittently on unrelated PRs) than 
with a code-level regression in this PR's diff - but I want to be upfront that 
I can't fully rule your diff out from static/log analysis alone, since I don't 
have a `dev`-branch control run of this exact test on this exact runner to 
compare against.
   
   Concrete next step I'd suggest, cheapest first: temporarily bump the 
wait-strategy timeout for just this one IT (or attach a log consumer from 
container start rather than only from the wait strategy) and do one more fork 
run. If the container eventually does emit a startup banner and register 
message, just later than 63s, that confirms plain scheduling/resource 
contention and a rerun or a slightly longer timeout is all that's needed - no 
diff change required. If it stays completely silent even with a much longer 
window, that would point at something more structural (networking/image) and 
would be worth flagging separately from this PR's dependency change, since your 
code path doesn't do anything at construction time that should cause total 
silence.
   
   This is not a known, previously-documented flake signature for this specific 
test on my side, so I can't just tell you "yes, that's a known issue, ignore 
it" - but based on the above I also don't think it's your fix that's causing 
it. I'd treat one more attempt with better diagnostics as the fastest way to 
get a real answer rather than guessing further from logs alone.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to