xinnyuli commented on issue #12382:
URL: https://github.com/apache/seatunnel/issues/12382#issuecomment-5829751870

   > Thanks [@xinnyuli](https://github.com/xinnyuli). I checked current `dev` 
at `1325a44b2f4013198fdb2c12f9df345de1cde4bc` and the related merged fix 
[#11864](https://github.com/apache/seatunnel/pull/11864). That fix changes the 
committed-offset boundary (`last_commit_lsn` and the non-exactly-once 
comparison); it does not establish that the post-savepoint path in this report 
cannot lose a record after the WAL task hands it to the reader.
   > 
   > There is no open PR covering this exact failure signature. Your 
investigation scope is the right next step, with one important constraint: 
please keep temporary tracing out of a proposed production fix. First make the 
Zeta savepoint -> restore -> post-reattachment insert sequence deterministic, 
then use the test-only evidence to distinguish these boundaries:
   > 
   > 1. `PostgresWalFetchTask` decoding and the record handed to the shared 
incremental fetcher;
   > 2. reader emission and checkpoint/offset advancement; and
   > 3. receipt and commit by the JDBC sink after restore.
   > 
   > The regression must assert both that the inserted row reaches the sink and 
that the reported/committed LSN does not advance past a record that was not 
delivered. If the repro isolates a separate boundary from 
[#11864](https://github.com/apache/seatunnel/pull/11864), please open one 
focused PR with that regression and the smallest corresponding fix. We should 
not add `help wanted` or choose a fix direction until that boundary is 
demonstrated.
   
   Thanks for the guidance. I’ve exercised the real Zeta savepoint → restore → 
post-reattachment insert flow locally and correlated the restored offsets with 
id=15 and its physical WAL record.
   The captured runs passed. Additional WAL records separated the saved offset 
from the INSERT, so they did not reproduce the suspected equal-LSN condition. 
These latest captures used an experimental resume implementation and are not an 
old-code/new-code comparison.
   I haven’t yet isolated a loss between decoding, reader emission, and sink 
commit, or demonstrated offsets advancing past an undelivered record. I 
therefore don’t have enough evidence to propose a fix for this issue.
   Could you or the reporter share the exact failing CI job link, commit SHA, 
and full engine/PostgreSQL logs? I’d like to compare the original failure with 
these captures before choosing a fix direction. Thanks!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to