SEZ9 commented on issue #12382:
URL: https://github.com/apache/seatunnel/issues/12382#issuecomment-5923595827

   Classification: A / Connector-V2 PostgreSQL CDC restore-delivery integrity.
   
   @xinnyuli thanks for the correction and the upstream context. Agreed on the 
framing: the locator logs and the +312-byte movement give us a concrete 
ambiguous-boundary hypothesis, but the link to the `id=15` row is still 
inferred, not shown. Let's keep that distinction explicit until the real-run 
comparison is in.
   
   On the upstream reference: the dbz#1554 / a51931d92 change (adding 
`lastProcessedMessageType` to the offset) is about several messages sharing one 
LSN under COPY, so as you say it doesn't automatically cover "stored commit-end 
LSN == start LSN of the next record". Whether it covers our case is exactly 
what your proposed measurement should settle, so please don't treat it as a 
confirmed fix path yet.
   
   Your next step is the right one. Concrete asks for the same two revisions, 
real Zeta savepoint -> restore path, uninstrumented where possible:
   1. In a failing run, record `pg_current_wal_lsn()` immediately before and 
after the `id=15` insert so we have the WAL range of that transaction.
   2. In the same run, capture the stored offset at savepoint and the LSN the 
locator logs as "already processed".
   3. Post the three values side by side and state whether the equality holds. 
Keep the two-revision/load matrix and the reached/not-reached accounting intact 
so we can see which runs the equality actually explains.
   
   If the equality holds, a deterministic regression test built on the real 
trigger order (stored commit-end LSN == next record's LSN) is what we want; a 
synthetic same-LSN unit test alone would only show the locator behavior, not 
that it explains this issue, as you noted.
   
   On the fix directions: a Debezium upgrade to 3.7 isn't viable while 
SeaTunnel pins 1.9.8.Final and supports Java 8/11, so that's off the table for 
now. Between porting a guard (offset format change, needs a compatibility path 
for existing savepoints) and a narrower restore-side workaround, I'd rather not 
pick until the measurement confirms the trigger. Also note that #12454 was 
flagged in this thread as touching the same recovery boundary 
(message-type-aware offset handling with existing savepoints still readable); 
please have a look and, if it overlaps, attach your real-run comparison there 
rather than opening a separate production fix. If it doesn't overlap, say so 
here and we'll decide the direction.
   
   No PR needed from you yet; the LSN comparison is the blocking item.
   
   <!-- streview-comment:1441 -->


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to