dev-donghwan commented on PR #10323:
URL: https://github.com/apache/paimon/pull/10323#issuecomment-6013126412

   You are right that this is an extreme scenario, and it does not happen 
often. We did hit it once in production, but it was a coincidence of a 
configuration issue on our side and an infrastructure issue at the same time. 
We have fixed that configuration, and I agree it is unlikely to happen again 
now. Here is what happened, with simplified numbers 
(`snapshot.num-retained.max=20`, checkpoints every minute):
   
   1. The latest snapshot and the LATEST hint are both 100.
   2. When committing 101, snapshot-101 was written to the storage, but a 
transient infrastructure issue made the response come back as a failure. The 
retry found snapshot-101 already there and finished the commit as successful, 
but that path did not update the hint.
   3. While the issue lasted, the same happened up to 121, so snapshots kept 
piling up while the hint stayed at 100.
   4. Expiration deleted 100 and 101. From then on, `findLatest` returned the 
deleted 100, so reads and commits all failed. Only a commit can rewrite the 
hint, and it failed first, so the table did not recover even after the issue 
was gone; we fixed LATEST by hand.
   
   The retry path in step 2 is fixed by #10324. But as long as the snapshot and 
the hint are written separately, an infrastructure issue between them cannot be 
ruled out, and when expiration then deletes the snapshot its own hint points 
to, a transient issue becomes an outage that needs manual repair. This PR only 
aims to prevent that last step.
   
   If you think this case is not worth the change, I will follow your judgment 
and close the PR.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to