xinnyuli commented on issue #12382:
URL: https://github.com/apache/seatunnel/issues/12382#issuecomment-5891344351

   > [@xinnyuli](https://github.com/xinnyuli), thanks for taking this on.
   > 
   > The comparison boundary is the failing scheduled run 
(https://github.com/apache/seatunnel/actions/runs/35244548330/job/105282609961, 
`dev` commit `c7304ace6e18d350314e92480df1fd3c0962f1f2`) versus current `dev` 
including `5af8d789aa9ae3d94df4a9cc0f03cee3c0a6d0e6`.
   > 
   > When you have results, could you post back with:
   > 
   > 1. the reproduction rate of the `id=15` restore assertion at each of those 
two revisions, with the exact revision tested recorded; and
   > 2. if it does reproduce, which of the three boundaries the row goes 
missing at (`PostgresWalFetchTask` handoff, reader emission/checkpoint-offset 
advancement, or JDBC-sink receipt/commit), using test-only, correlation-safe 
evidence.
   > 
   > If it does not reproduce with the real Zeta savepoint -> restore -> 
post-reattachment sequence at either revision, please say so explicitly. That 
result will decide the next step; until then, no production-fix PR or 
production tracing is needed.
   
   Results for the id=15 restore assertion, as requested.
   
   **Setup.** 
`PostgresCDCIT#testPostgresCdcSnapshotOnlyAndCommittedOffsetStartupModes`, real 
Zeta savepoint -> restore -> post-reattachment INSERT (id=15), zeta container 
only (`RUN_ALL_CONTAINER=false`, `RUN_ZETA_CONTAINER=true`), GitHub-hosted 
`ubuntu-latest`, JDK 11 (the failing job was the Java 11 matrix entry), `-Pci`, 
`-Xmx4096m`. One run per parallel job, fresh container each time. Test-only: an 
observation patch that adds log lines to the IT plus a javaagent loaded into 
the test's server JVM. No production code changed; no experimental resume 
applied.
   
   **Revisions (recorded by the workflow):**
   - failing revision: `c7304ace6e18d350314e92480df1fd3c0962f1f2`
   - current dev: `4c874e2a4061aea9d5db65e74edeb211b498fe27` (verified to 
contain #11864, `5af8d789aa9ae3d94df4a9cc0f03cee3c0a6d0e6`)
   
   **Result, 30 runs each:**
   
   | revision | reached the id=15 check | id=15 present in sink | row missing | 
did not reach the check |
   |---|---|---|---|---|
   | `c7304ace` | 29 | 29 | 0 | 1 (setup timeout) |
   | `4c874e2a` (dev) | 29 | 29 | 0 | 1 (setup timeout) |
   
   **I could not reproduce the loss at either revision.** 0/29 at each; the 
exact one-sided 95% upper bound on the per-run failure rate is 9.8% for each 
(about 5% pooled). Since nothing failed, I have no evidence for which of the 
three boundaries would lose the row, and I am not claiming one.
   
   The one non-passing run per revision failed earlier, in the setup phase (the 
wait for the slot's committed LSN to reach the post-seed LSN, ~line 489), 
before the restore and before id=15 was inserted. It is a different failure 
from this issue, so I counted it separately and did not count it as a pass. I 
also saw the same setup timeout once in 5 local runs.
   
   Limits: the tracing agent adds timing overhead, and hosted runners differ 
from the original one, so this bounds this setup only. Next I plan to (1) 
re-run the same matrix with CPU/memory/IO contention (stress-ng) to see whether 
load changes the rate, and (2) run an uninstrumented control. I will post those 
results here. Run details and logs: 
<https://github.com/xinnyuli/seatunnel-12382-repro/actions/runs/36571981978>.
   
   No production-fix PR or production tracing from my side.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to