JTaky opened a new pull request, #10156: URL: https://github.com/apache/paimon/pull/10156
Paimon's filesystem catalog over an object store (e.g. S3) needs an external lock: RenamingSnapshotCommit.commit is only safe under lock.runWithLock(...) because rename is not atomic on an object store, so two writers picking the same snapshot id can both "succeed" and one snapshot is silently lost. Paimon ships Jdbc and Hive lock implementations for this; this adds a ZooKeeper one, selected via lock.type=zookeeper. Built on a Curator InterProcessMutex (ephemeral-node acquire/run/release, matching Paimon's own lock contract), with two additional safeguards beyond a bare mutex: - A session-loss epoch, bumped by a ConnectionStateListener on SUSPENDED/LOST/RECONNECTED, checked after the callable returns. - A synchronous checkExists() round trip against the held znode after the callable returns, which is what actually closes the race window isAcquiredInThisProcess() alone cannot: that flag is local bookkeeping Curator never updates on connection loss, so it stays true straight through a genuine session expiry. Both failure paths throw rather than report success, turning a potential silent lost snapshot into a visible, retried commit failure, with stable log markers (paimon_lock_lost, paimon_lock_acquire_timeout, paimon_lock_acquire_error, paimon_lock_release_failed) for monitoring. One CuratorFramework is shared per (quorum, root) per JVM, since a process can hold many (committer, table) locks and each must not open its own connection. ### Purpose ### Tests -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
