davidzollo opened a new pull request, #12103:
URL: https://github.com/apache/seatunnel/pull/12103
## Purpose
Adds E2E coverage documenting a verified, currently-open architectural gap:
under `ScheduleStrategy.WAIT`, `CoordinatorService#pendingJobSchedule`
retries a permanently-unschedulable head-of-queue job every ~3s forever,
uncapped, and never dequeues it. Since the single-threaded scheduler always
re-peeks the FIFO head, this starves every job submitted after it
indefinitely, even ones that are trivially schedulable with
currently-available resources.
This is not a regression test for a past fix — it's a test that documents
today's actual behavior precisely, so any future change to it is a
deliberate, visible diff rather than a silent one.
## Why this needed fresh verification, not just trusting an older analysis
This same area of code was previously touched by #11653 (already merged),
which added `reservePendingJobInfo`/`releasePendingJobInfo` (an
epoch-based reservation mechanism) and a predicate-based
`PeekBlockingQueue#peekBlocking`, specifically to stop a stale scheduler
generation from double-dispatching a job after a master failover.
I traced this mechanism against current `CoordinatorService.java` and
`PeekBlockingQueue.java` directly rather than assuming the older finding
still held: that reservation set only excludes a job id already reserved
by *another in-flight scheduler thread across overlapping generations*
during a master flip — it is not, and was never meant to be, a "this job
already failed its resource check" marker. In steady state (single
scheduler generation, no failover in progress), `schedulingPendingJobIds`
is empty at every peek, so the head job always matches and is returned
again. The head-of-line blocking is real and fully reproducible on
current `dev`.
## What the test does
`SplitClusterPendingJobLifecycleFailoverIT#testPermanentlyStuckWaitJobBlocksSchedulableJobBehindIt`:
1. Submits a job with unsatisfiable parallelism (999) under
`ScheduleStrategy.WAIT` — it can never acquire enough slots.
2. Submits a second, easily-schedulable job right behind it.
3. Confirms the second job stays in `PENDING` for a sustained window
(proving it isn't merely slow — it's genuinely blocked).
4. Cancels the stuck first job and confirms the second job reaches
`RUNNING` promptly, with no other change to cluster resources —
proving it was schedulable the entire time.
## Test plan
- New test method + one new minimal test-resource config; no production
code changed.
- `./mvnw spotless:apply` on the affected module — succeeded.
- `./mvnw install -DskipTests` (module + dependency chain,
`seatunnel-engine-ui` excluded since unmodified) — confirmed genuine
compilation via the actual `.class` file present under
`target/test-classes` with `javap -p` showing the new test method and
constant in the compiled bytecode (not run with
`-Dmaven.test.skip=true`, which skips test compilation entirely rather
than just execution).
- Full test execution is left to CI per this repository's E2E conventions
(Hazelcast-cluster-backed, not run in this sandbox).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]