DanielLeens commented on PR #11727:
URL: https://github.com/apache/seatunnel/pull/11727#issuecomment-5601140161

   @abdessalems thanks for filing #12224 for the `CooperativeTaskWorker` gap 
and linking it here alongside #12164 — that was the one remaining ask on my 
side, so I have no further source-side blocker on this PR. Nice work re-running 
just the two unrelated failed jobs instead of the whole workflow too.
   
   @SEZ9 — re: the truncation, I re-pulled my own comment's stored body 
directly via the API (not the web renderer) and it is not cut off there either: 
after `executionContexts.remove(taskGroupLocation, ownedContext)` — it 
continues "a compare-and-remove keyed on the exact context object this tracker 
owns. That call is safe against *any* current map value: if the map holds 
something other than `ownedContext` (a leaked entry, a newer generation's 
context, or nothing at all), the remove simply returns `false` and 
`finishOwnedResources` takes the "stale" branch (`:1571-1576`), preserving 
whatever is there." So — same as the rendering artifact you hit on my end 
earlier in the thread — this looks like a display-side truncation on your end, 
not a gap in what's actually stored.
   
   To answer your question directly, since it matters more than the transport 
glitch: **the `remove(key, value)` path tolerates a leaked entry being 
present.** It's a compare-and-remove, not an unconditional remove, so an 
unexpected value (leaked entry, a different generation's context, or absence) 
all fall through to the same "stale, do nothing to this entry" branch rather 
than corrupting state or throwing. The ownership logic's correctness does not 
depend on the leaked entry being absent — it was written to tolerate exactly 
that case.
   
   What the leak does break is upstream, not in this PR's own logic: a leaked 
`executionContexts` entry from a submit-time failure (#12164) means no 
`TaskGroupExecutionTracker` for that generation ever runs to call 
`finishOwnedResources()`, and `deployTask`'s outer guard treats any live entry 
as "already active," so no redeploy can even create the next generation to own 
that slot. This PR's ownership path is correct in the presence of the leak, 
just starved by it — which is why the leaked-entry point stays folded into the 
Blocker 1 tracking split (independent bugs that happen to compound in one 
failure window) rather than needing its own rollback in this PR.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to