xinnyuli commented on issue #12382:
URL: https://github.com/apache/seatunnel/issues/12382#issuecomment-5893897790

   > Classification: A / Connector-V2 PostgreSQL CDC restore-delivery integrity.
   > 
   > Thank you for the disciplined comparison. Thirty runs on each recorded 
revision, with the real Zeta savepoint -> restore -> post-reattachment insert 
and no lost `id=15` result in either 29-run reached sample, materially weakens 
the current loss hypothesis. The result is not a confirmation that the original 
report is fixed: the observed zero-event upper bound is still 9.8% per 
revision, the original scheduled failure remains unexplained, and the agent 
instrumentation can change timing.
   > 
   > Please keep this as an investigation, not a production-fix path. The next 
useful evidence is an uninstrumented control on the same two revisions, then a 
separately recorded controlled-load matrix if it is needed. For every run, 
retain the reached/not-reached count and keep the pre-restore 
slot/committed-LSN setup timeout distinct from the delivery assertion. Only if 
the row loss becomes reproducible should test-only correlation trace the WAL 
handoff, reader emission/checkpoint advancement, and JDBC-sink receipt. No 
production tracing or fix PR is justified yet.
   > 
   > No label or assignment change in this pass.
   
   Thanks! Both follow-ups you asked for are done; everything is on 
GitHub-hosted ubuntu-latest, Java 11, same two revisions (baseline c7304ace, 
dev 4c874e2a which contains #11864), test-only, no production code changed.
   
   Without load, as before: 0 lost rows in 29/29 reached runs per revision.
   
   Controlled load (stress-ng running while the IT executes), recorded 
separately:
   
   | revision | instrumentation | runs | reached id=15 check | row missing | 
setup timeout (before the check) |
   |---|---|---|---|---|---|
   | baseline | none (untouched upstream test) | 30 | 30 | 2 | 0 |
   | dev | none | 30 | 29 | 1 | 1 |
   | baseline | tracing agent | 30 | 29 | 4 | 1 |
   | dev | tracing agent | 30 | 30 | 3 | 0 |
   
   "Row missing" is only the delivery assertion (`expected: <1> but was: <0>`, 
id=15 not in the sink within 3 minutes). The setup timeout on the committed-LSN 
wait is counted separately and is not in the denominator. The uninstrumented 
failures are the same assertion, so the failure reproduces without any tracing; 
the difference between the two groups is not statistically meaningful at this 
sample size, so I can't say whether tracing changes the rate. The original 
scheduled failure is still unexplained and a hosted runner under artificial 
load is not the same as the original job.
   
   Since row loss is now reproducible under load, I looked at the test-only 
trace in the 7 missing-row runs: no id=15 record reaches the Debezium receiver, 
while the same trace follows the row through receiver -> emitter -> sink write 
in all 52 passing runs. I have not yet checked whether Postgres does not send 
it or Debezium drops it, and I don't want to add more tracing unless you think 
that is worthwhile.
   
   I'm not proposing a production fix or production tracing. Happy to stop 
here, or to add one more test-only hook below the receiver if that is useful to 
you.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to