xinnyuli commented on issue #12382: URL: https://github.com/apache/seatunnel/issues/12382#issuecomment-5891344351
> [@xinnyuli](https://github.com/xinnyuli), thanks for taking this on. > > The comparison boundary is the failing scheduled run (https://github.com/apache/seatunnel/actions/runs/35244548330/job/105282609961, `dev` commit `c7304ace6e18d350314e92480df1fd3c0962f1f2`) versus current `dev` including `5af8d789aa9ae3d94df4a9cc0f03cee3c0a6d0e6`. > > When you have results, could you post back with: > > 1. the reproduction rate of the `id=15` restore assertion at each of those two revisions, with the exact revision tested recorded; and > 2. if it does reproduce, which of the three boundaries the row goes missing at (`PostgresWalFetchTask` handoff, reader emission/checkpoint-offset advancement, or JDBC-sink receipt/commit), using test-only, correlation-safe evidence. > > If it does not reproduce with the real Zeta savepoint -> restore -> post-reattachment sequence at either revision, please say so explicitly. That result will decide the next step; until then, no production-fix PR or production tracing is needed. Results for the id=15 restore assertion, as requested. **Setup.** `PostgresCDCIT#testPostgresCdcSnapshotOnlyAndCommittedOffsetStartupModes`, real Zeta savepoint -> restore -> post-reattachment INSERT (id=15), zeta container only (`RUN_ALL_CONTAINER=false`, `RUN_ZETA_CONTAINER=true`), GitHub-hosted `ubuntu-latest`, JDK 11 (the failing job was the Java 11 matrix entry), `-Pci`, `-Xmx4096m`. One run per parallel job, fresh container each time. Test-only: an observation patch that adds log lines to the IT plus a javaagent loaded into the test's server JVM. No production code changed; no experimental resume applied. **Revisions (recorded by the workflow):** - failing revision: `c7304ace6e18d350314e92480df1fd3c0962f1f2` - current dev: `4c874e2a4061aea9d5db65e74edeb211b498fe27` (verified to contain #11864, `5af8d789aa9ae3d94df4a9cc0f03cee3c0a6d0e6`) **Result, 30 runs each:** | revision | reached the id=15 check | id=15 present in sink | row missing | did not reach the check | |---|---|---|---|---| | `c7304ace` | 29 | 29 | 0 | 1 (setup timeout) | | `4c874e2a` (dev) | 29 | 29 | 0 | 1 (setup timeout) | **I could not reproduce the loss at either revision.** 0/29 at each; the exact one-sided 95% upper bound on the per-run failure rate is 9.8% for each (about 5% pooled). Since nothing failed, I have no evidence for which of the three boundaries would lose the row, and I am not claiming one. The one non-passing run per revision failed earlier, in the setup phase (the wait for the slot's committed LSN to reach the post-seed LSN, ~line 489), before the restore and before id=15 was inserted. It is a different failure from this issue, so I counted it separately and did not count it as a pass. I also saw the same setup timeout once in 5 local runs. Limits: the tracing agent adds timing overhead, and hosted runners differ from the original one, so this bounds this setup only. Next I plan to (1) re-run the same matrix with CPU/memory/IO contention (stress-ng) to see whether load changes the rate, and (2) run an uninstrumented control. I will post those results here. Run details and logs: <https://github.com/xinnyuli/seatunnel-12382-repro/actions/runs/36571981978>. No production-fix PR or production tracing from my side. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
