SEZ9 commented on issue #12382: URL: https://github.com/apache/seatunnel/issues/12382#issuecomment-5923595827
Classification: A / Connector-V2 PostgreSQL CDC restore-delivery integrity. @xinnyuli thanks for the correction and the upstream context. Agreed on the framing: the locator logs and the +312-byte movement give us a concrete ambiguous-boundary hypothesis, but the link to the `id=15` row is still inferred, not shown. Let's keep that distinction explicit until the real-run comparison is in. On the upstream reference: the dbz#1554 / a51931d92 change (adding `lastProcessedMessageType` to the offset) is about several messages sharing one LSN under COPY, so as you say it doesn't automatically cover "stored commit-end LSN == start LSN of the next record". Whether it covers our case is exactly what your proposed measurement should settle, so please don't treat it as a confirmed fix path yet. Your next step is the right one. Concrete asks for the same two revisions, real Zeta savepoint -> restore path, uninstrumented where possible: 1. In a failing run, record `pg_current_wal_lsn()` immediately before and after the `id=15` insert so we have the WAL range of that transaction. 2. In the same run, capture the stored offset at savepoint and the LSN the locator logs as "already processed". 3. Post the three values side by side and state whether the equality holds. Keep the two-revision/load matrix and the reached/not-reached accounting intact so we can see which runs the equality actually explains. If the equality holds, a deterministic regression test built on the real trigger order (stored commit-end LSN == next record's LSN) is what we want; a synthetic same-LSN unit test alone would only show the locator behavior, not that it explains this issue, as you noted. On the fix directions: a Debezium upgrade to 3.7 isn't viable while SeaTunnel pins 1.9.8.Final and supports Java 8/11, so that's off the table for now. Between porting a guard (offset format change, needs a compatibility path for existing savepoints) and a narrower restore-side workaround, I'd rather not pick until the measurement confirms the trigger. Also note that #12454 was flagged in this thread as touching the same recovery boundary (message-type-aware offset handling with existing savepoints still readable); please have a look and, if it overlaps, attach your real-run comparison there rather than opening a separate production fix. If it doesn't overlap, say so here and we'll decide the direction. No PR needed from you yet; the LSN comparison is the blocking item. <!-- streview-comment:1441 --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
