DanielLeens opened a new issue, #12125:
URL: https://github.com/apache/seatunnel/issues/12125

   ## Description
   
   This is an investigation issue, not a bug report. The code gap below is 
confirmed; the failure chain is a hypothesis that needs a merge-specific 
reproduction on `dev` before it is treated as a defect.
   
   Verified at `dev` commit `97d461bc0773399d632fd078735736ecd44f5f0b`:
   
   - `SeaTunnelServer.java:224` is `public void reset() {}`. `SeaTunnelServer 
implements ManagedService, MembershipAwareService, LiveOperationsTracker`; 
Hazelcast 5.1's `ClusterMergeTask` resets managed services during a cluster 
merge, so this is the callback that runs when a member rejoins after a split.
   - Nothing on that path touches `TaskExecutionService.executionContexts` 
(`TaskExecutionService.java:184`) or `finishedExecutionContexts` (`:191`).
   - The idempotent redeploy branch from #10567 
(`TaskExecutionService.java:506-525`) returns `TaskDeployState.success()` when 
`executionContexts.containsKey(location)` and compares no execution epoch, so 
it cannot distinguish a genuinely active context from one that outlived a 
membership reset.
   
   Not established: that a frozen worker merges on the losing side, retains an 
invalid execution, later receives the same `TaskGroupLocation`, is adopted by 
the new master, and then stops checkpointing. Normal takeover also runs 
`PhysicalVertex.initStateFuture` and `CheckpointCoordinator.restoreCoordinator` 
(`:673`), so an existing context is not by itself evidence that takeover fails. 
A master-only rolling restart (#10568, fixed by #10567) and a losing-side 
cluster merge are different paths. #8566 and #9061 predate #10567 and cannot 
show a separate post-fix root cause. A downstream fork reproduced freeze -> 
evict -> rejoin -> `already exists` on a codebase without #10567; on `dev` the 
consequence has not been reproduced.
   
   ## Expected outcome
   
   1. Reproduction on `dev`: `SIGSTOP` one worker past the heartbeat timeout, 
confirm eviction, `SIGCONT`, capture the merge and the `reset()` invocation, 
then kill the active master. Record member identities, task locations, slot 
ownership, and whether checkpoints keep advancing.
   2. If the gap reproduces: implement `reset()` so that it cancels locally 
owned active task groups and lets the existing completion path release each 
resource exactly once (`taskDone` moves the context at `:1459-1460`; 
`recycleClassLoader` releases loaders at `:1507-1511`), rather than clearing 
the maps directly; make it idempotent across repeated reset and failover; add a 
regression test with the same shape.
   3. Related: #11437 / #11458 (slot-side reconciliation after failover), 
#10568 / #10567, #10675 / #10692.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to