davidzollo opened a new pull request, #12103:
URL: https://github.com/apache/seatunnel/pull/12103

   ## Purpose
   
   Adds E2E coverage documenting a verified, currently-open architectural gap:
   under `ScheduleStrategy.WAIT`, `CoordinatorService#pendingJobSchedule`
   retries a permanently-unschedulable head-of-queue job every ~3s forever,
   uncapped, and never dequeues it. Since the single-threaded scheduler always
   re-peeks the FIFO head, this starves every job submitted after it
   indefinitely, even ones that are trivially schedulable with
   currently-available resources.
   
   This is not a regression test for a past fix — it's a test that documents
   today's actual behavior precisely, so any future change to it is a
   deliberate, visible diff rather than a silent one.
   
   ## Why this needed fresh verification, not just trusting an older analysis
   
   This same area of code was previously touched by #11653 (already merged),
   which added `reservePendingJobInfo`/`releasePendingJobInfo` (an
   epoch-based reservation mechanism) and a predicate-based
   `PeekBlockingQueue#peekBlocking`, specifically to stop a stale scheduler
   generation from double-dispatching a job after a master failover.
   
   I traced this mechanism against current `CoordinatorService.java` and
   `PeekBlockingQueue.java` directly rather than assuming the older finding
   still held: that reservation set only excludes a job id already reserved
   by *another in-flight scheduler thread across overlapping generations*
   during a master flip — it is not, and was never meant to be, a "this job
   already failed its resource check" marker. In steady state (single
   scheduler generation, no failover in progress), `schedulingPendingJobIds`
   is empty at every peek, so the head job always matches and is returned
   again. The head-of-line blocking is real and fully reproducible on
   current `dev`.
   
   ## What the test does
   
   
`SplitClusterPendingJobLifecycleFailoverIT#testPermanentlyStuckWaitJobBlocksSchedulableJobBehindIt`:
   
   1. Submits a job with unsatisfiable parallelism (999) under
      `ScheduleStrategy.WAIT` — it can never acquire enough slots.
   2. Submits a second, easily-schedulable job right behind it.
   3. Confirms the second job stays in `PENDING` for a sustained window
      (proving it isn't merely slow — it's genuinely blocked).
   4. Cancels the stuck first job and confirms the second job reaches
      `RUNNING` promptly, with no other change to cluster resources —
      proving it was schedulable the entire time.
   
   ## Test plan
   
   - New test method + one new minimal test-resource config; no production
     code changed.
   - `./mvnw spotless:apply` on the affected module — succeeded.
   - `./mvnw install -DskipTests` (module + dependency chain,
     `seatunnel-engine-ui` excluded since unmodified) — confirmed genuine
     compilation via the actual `.class` file present under
     `target/test-classes` with `javap -p` showing the new test method and
     constant in the compiled bytecode (not run with
     `-Dmaven.test.skip=true`, which skips test compilation entirely rather
     than just execution).
   - Full test execution is left to CI per this repository's E2E conventions
     (Hazelcast-cluster-backed, not run in this sandbox).


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to