goutamadwant commented on issue #11735: URL: https://github.com/apache/seatunnel/issues/11735#issuecomment-5809848543
@DanielLeens @SEZ9 Thanks for spelling out the remaining boundaries. I checked each one against current `dev` (`7067a0984`) and updated the STIP body above and #11734 (EN and ZH, commit `e944f1e3`). The two copies now carry the same text. **1. Diagnostic attempt under active-master change** - The attempt key is the pipeline's `CREATED` timestamp. `dev` writes it only in the `SubPlan` constructor and `resetPipelineState()`, and a rebuilt `SubPlan` keeps it. - The key is copied at reset and deploy time and stored in history. It is never read back from `runningJobStateTimestampsIMap` later, which avoids the overwrite problem raised in August. - Every record carries its own key, so a lost acknowledgement, a retry or a late registration resolves to the same attempt. `resetPipelineState()` writes `max(now, previousCreated + 1)`, so keys stay ordered across masters. This value is returned by the `/job-info` diagnostics, so the change is stated in Compatibility and in the `rest-api-v2` item. - Worker reports carry the per-deployment `executionId`. After failover, surviving task groups resolve through placeholder deployments. - `canRestorePipeline()` and `job.retry.times` never read this state, and no restore step waits for it. The crash-window table now covers every pipeline state at failover, including the pending-queue window, where reports are already lost today. This replaces the earlier "advance must complete before the restored execution starts" rule with the reconciliation form Daniel mentioned as an alternative in August. Blocking restore on a diagnostic write would have changed restore availability. **2. Best-effort writes** - Capture only does a bounded copy and a non-blocking `offer` into a per-master queue. The queue is bounded by count, bytes and a per-job share. - A dedicated consumer thread runs the `EntryProcessor`s. It retries only Hazelcast-retryable failures, at most three times, and then drops the operation with a WARN that carries identifiers only. - Completion never touches task or pipeline state or re-enters failure/restore processing. - Finalization goes through the same queue, so the failures that led to the terminal state are applied before the entry becomes terminal. **3. Wire compatibility** - `TaskExecutionState` (`-108652017022658969L`) and `TaskDeployState` (`2646079648150626562L`) have no declared UID today. These are their generated values from `serialver` on JDK 8 and 11. - Both get exactly these values pinned before any field is added. - The new fields are nullable JDK types only, tested with byte fixtures in both directions plus a UID guard. - `JobInfo` and `JobCleanupRecord` are unchanged. Cleanup of the history map does not extend `JobCleanupRecord`. **4. Tests** The acceptance criteria list the tests for each boundary above: - crash windows: 9 and 10; - late writes after terminal cleanup: 20, 21 and 23; - known job with no records vs unknown or expired job: 25. While checking against `dev` I also fixed some gaps in my own text: - stale reports from a previous deployment; - restores without a task-group failure (checkpoint, resource, master recovery); - the history maps are now excluded from `map-store` persistence; - a finished-snapshot budget and memory formula; - the REST lookup order during cleanup and savepoint restarts; - a stricter URI grammar and fixed error bodies; - the redaction order and patterns. This is still a design proposal. Please take another look when you have a chance. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
