dev-donghwan commented on PR #10323: URL: https://github.com/apache/paimon/pull/10323#issuecomment-6013126412
You are right that this is an extreme scenario, and it does not happen often. We did hit it once in production, but it was a coincidence of a configuration issue on our side and an infrastructure issue at the same time. We have fixed that configuration, and I agree it is unlikely to happen again now. Here is what happened, with simplified numbers (`snapshot.num-retained.max=20`, checkpoints every minute): 1. The latest snapshot and the LATEST hint are both 100. 2. When committing 101, snapshot-101 was written to the storage, but a transient infrastructure issue made the response come back as a failure. The retry found snapshot-101 already there and finished the commit as successful, but that path did not update the hint. 3. While the issue lasted, the same happened up to 121, so snapshots kept piling up while the hint stayed at 100. 4. Expiration deleted 100 and 101. From then on, `findLatest` returned the deleted 100, so reads and commits all failed. Only a commit can rewrite the hint, and it failed first, so the table did not recover even after the issue was gone; we fixed LATEST by hand. The retry path in step 2 is fixed by #10324. But as long as the snapshot and the hint are written separately, an infrastructure issue between them cannot be ruled out, and when expiration then deletes the snapshot its own hint points to, a transient issue becomes an outage that needs manual repair. This PR only aims to prevent that last step. If you think this case is not worth the change, I will follow your judgment and close the PR. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
