JTaky opened a new pull request, #10156:
URL: https://github.com/apache/paimon/pull/10156

   Paimon's filesystem catalog over an object store (e.g. S3) needs an external 
lock: RenamingSnapshotCommit.commit is only safe under lock.runWithLock(...) 
because rename is not atomic on an object store, so two writers picking the 
same snapshot id can both "succeed" and one snapshot is silently lost. Paimon 
ships Jdbc and Hive lock implementations for this; this adds a ZooKeeper one, 
selected via lock.type=zookeeper.
   
   Built on a Curator InterProcessMutex (ephemeral-node acquire/run/release, 
matching Paimon's own lock contract), with two additional safeguards beyond a 
bare mutex:
   - A session-loss epoch, bumped by a ConnectionStateListener on 
SUSPENDED/LOST/RECONNECTED, checked after the callable returns.
   - A synchronous checkExists() round trip against the held znode after the 
callable returns, which is what actually closes the race window 
isAcquiredInThisProcess() alone cannot: that flag is local bookkeeping Curator 
never updates on connection loss, so it stays true straight through a genuine 
session expiry.
   
   Both failure paths throw rather than report success, turning a potential 
silent lost snapshot into a visible, retried commit failure, with stable log 
markers (paimon_lock_lost, paimon_lock_acquire_timeout, 
paimon_lock_acquire_error, paimon_lock_release_failed) for monitoring.
   
   One CuratorFramework is shared per (quorum, root) per JVM, since a process 
can hold many (committer, table) locks and each must not open its own 
connection.
   
   ### Purpose
   
   ### Tests
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to