xinnyuli commented on issue #12382: URL: https://github.com/apache/seatunnel/issues/12382#issuecomment-5829751870
> Thanks [@xinnyuli](https://github.com/xinnyuli). I checked current `dev` at `1325a44b2f4013198fdb2c12f9df345de1cde4bc` and the related merged fix [#11864](https://github.com/apache/seatunnel/pull/11864). That fix changes the committed-offset boundary (`last_commit_lsn` and the non-exactly-once comparison); it does not establish that the post-savepoint path in this report cannot lose a record after the WAL task hands it to the reader. > > There is no open PR covering this exact failure signature. Your investigation scope is the right next step, with one important constraint: please keep temporary tracing out of a proposed production fix. First make the Zeta savepoint -> restore -> post-reattachment insert sequence deterministic, then use the test-only evidence to distinguish these boundaries: > > 1. `PostgresWalFetchTask` decoding and the record handed to the shared incremental fetcher; > 2. reader emission and checkpoint/offset advancement; and > 3. receipt and commit by the JDBC sink after restore. > > The regression must assert both that the inserted row reaches the sink and that the reported/committed LSN does not advance past a record that was not delivered. If the repro isolates a separate boundary from [#11864](https://github.com/apache/seatunnel/pull/11864), please open one focused PR with that regression and the smallest corresponding fix. We should not add `help wanted` or choose a fix direction until that boundary is demonstrated. Thanks for the guidance. I’ve exercised the real Zeta savepoint → restore → post-reattachment insert flow locally and correlated the restored offsets with id=15 and its physical WAL record. The captured runs passed. Additional WAL records separated the saved offset from the INSERT, so they did not reproduce the suspected equal-LSN condition. These latest captures used an experimental resume implementation and are not an old-code/new-code comparison. I haven’t yet isolated a loss between decoding, reader emission, and sink commit, or demonstrated offsets advancing past an undelivered record. I therefore don’t have enough evidence to propose a fix for this issue. Could you or the reporter share the exact failing CI job link, commit SHA, and full engine/PostgreSQL logs? I’d like to compare the original failure with these captures before choosing a fix direction. Thanks! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
