xinnyuli commented on issue #12382: URL: https://github.com/apache/seatunnel/issues/12382#issuecomment-5893897790
> Classification: A / Connector-V2 PostgreSQL CDC restore-delivery integrity. > > Thank you for the disciplined comparison. Thirty runs on each recorded revision, with the real Zeta savepoint -> restore -> post-reattachment insert and no lost `id=15` result in either 29-run reached sample, materially weakens the current loss hypothesis. The result is not a confirmation that the original report is fixed: the observed zero-event upper bound is still 9.8% per revision, the original scheduled failure remains unexplained, and the agent instrumentation can change timing. > > Please keep this as an investigation, not a production-fix path. The next useful evidence is an uninstrumented control on the same two revisions, then a separately recorded controlled-load matrix if it is needed. For every run, retain the reached/not-reached count and keep the pre-restore slot/committed-LSN setup timeout distinct from the delivery assertion. Only if the row loss becomes reproducible should test-only correlation trace the WAL handoff, reader emission/checkpoint advancement, and JDBC-sink receipt. No production tracing or fix PR is justified yet. > > No label or assignment change in this pass. Thanks! Both follow-ups you asked for are done; everything is on GitHub-hosted ubuntu-latest, Java 11, same two revisions (baseline c7304ace, dev 4c874e2a which contains #11864), test-only, no production code changed. Without load, as before: 0 lost rows in 29/29 reached runs per revision. Controlled load (stress-ng running while the IT executes), recorded separately: | revision | instrumentation | runs | reached id=15 check | row missing | setup timeout (before the check) | |---|---|---|---|---|---| | baseline | none (untouched upstream test) | 30 | 30 | 2 | 0 | | dev | none | 30 | 29 | 1 | 1 | | baseline | tracing agent | 30 | 29 | 4 | 1 | | dev | tracing agent | 30 | 30 | 3 | 0 | "Row missing" is only the delivery assertion (`expected: <1> but was: <0>`, id=15 not in the sink within 3 minutes). The setup timeout on the committed-LSN wait is counted separately and is not in the denominator. The uninstrumented failures are the same assertion, so the failure reproduces without any tracing; the difference between the two groups is not statistically meaningful at this sample size, so I can't say whether tracing changes the rate. The original scheduled failure is still unexplained and a hosted runner under artificial load is not the same as the original job. Since row loss is now reproducible under load, I looked at the test-only trace in the 7 missing-row runs: no id=15 record reaches the Debezium receiver, while the same trace follows the row through receiver -> emitter -> sink write in all 52 passing runs. I have not yet checked whether Postgres does not send it or Debezium drops it, and I don't want to add more tracing unless you think that is worthwhile. I'm not proposing a production fix or production tracing. Happy to stop here, or to add one more test-only hook below the receiver if that is useful to you. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
