zhuxiangyi commented on PR #10105: URL: https://github.com/apache/paimon/pull/10105#issuecomment-5810032793
Thanks — all four reproduced first as failing tests, then fixed in b3dcfca (rebased on master): 1. **Explicit commit user across a new checkpoint**: removed `write.stream.commit-user` (new in this PR, unreleased). Batch ids are only unique within one checkpoint, so the identity is always the query id and not configurable; `commit.user-prefix` can still name it. The docs no longer recommend pinning it. 2. **Replay after expiration**: before each commit the sink records under the checkpoint location which snapshot was the latest. On a replay with no snapshot of the commit user left, if expiration has gone past that snapshot, the query fails naming the marker instead of committing again; otherwise the lookup is exact (expiration removes the oldest first). Documented as a retention bound. 3. **Compaction before the retry**: the retry follows the superseded files through the snapshot branch's compactions since the overwrite and clears their outputs; if a compaction merged them with later rows, it fails instead of reporting success. 4. **Overwrite-time snapshot expired**: the retry now fails explicitly when the snapshot branch no longer retains a snapshot from before the overwrite (and only skips when the branch had none then). Regression tests for each, mutation-checked. I also added a replay test for every kind of table the sink writes to (dynamic bucket, cross partition, postpone with real buckets, DV, lookup changelog, row tracking), a batch failing before its commit, and a query upgraded from the old per-batch commits; sink suites pass on Spark 3.2 / 3.4 / 3.5 / 4.0 / 4.1. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
