DanielLeens commented on PR #11727: URL: https://github.com/apache/seatunnel/pull/11727#issuecomment-5601140161
@abdessalems thanks for filing #12224 for the `CooperativeTaskWorker` gap and linking it here alongside #12164 — that was the one remaining ask on my side, so I have no further source-side blocker on this PR. Nice work re-running just the two unrelated failed jobs instead of the whole workflow too. @SEZ9 — re: the truncation, I re-pulled my own comment's stored body directly via the API (not the web renderer) and it is not cut off there either: after `executionContexts.remove(taskGroupLocation, ownedContext)` — it continues "a compare-and-remove keyed on the exact context object this tracker owns. That call is safe against *any* current map value: if the map holds something other than `ownedContext` (a leaked entry, a newer generation's context, or nothing at all), the remove simply returns `false` and `finishOwnedResources` takes the "stale" branch (`:1571-1576`), preserving whatever is there." So — same as the rendering artifact you hit on my end earlier in the thread — this looks like a display-side truncation on your end, not a gap in what's actually stored. To answer your question directly, since it matters more than the transport glitch: **the `remove(key, value)` path tolerates a leaked entry being present.** It's a compare-and-remove, not an unconditional remove, so an unexpected value (leaked entry, a different generation's context, or absence) all fall through to the same "stale, do nothing to this entry" branch rather than corrupting state or throwing. The ownership logic's correctness does not depend on the leaked entry being absent — it was written to tolerate exactly that case. What the leak does break is upstream, not in this PR's own logic: a leaked `executionContexts` entry from a submit-time failure (#12164) means no `TaskGroupExecutionTracker` for that generation ever runs to call `finishOwnedResources()`, and `deployTask`'s outer guard treats any live entry as "already active," so no redeploy can even create the next generation to own that slot. This PR's ownership path is correct in the presence of the leak, just starved by it — which is why the leaked-entry point stays folded into the Blocker 1 tracking split (independent bugs that happen to compound in one failure window) rather than needing its own rollback in this PR. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
