Hi All,

My impression from this thread is that it might be time for a call to
discuss all the related issues and try to arrive at a shared implementation
plan.

I know meetings are not ideal, but in this case it might be beneficial as a
means for achieving a common understanding of the set of problems and
priorities related to this thread. A dedicated metrics meeting worked well
from my POV.

Allocating a time slot in the next community meeting might be an option,
although I think we might need the full hour in this case.

WDYT?

Thanks,
Dmitri.

On Mon, Aug 17, 2026 at 6:39 PM Prithvi S <[email protected]>
wrote:

> Hi all,
>
> I rewrote https://github.com/apache/polaris/pull/5035, following the
> discussion to focus only on the SPI foundation.
>
> What is in the PR now:
>
>    - A written manager-level consistency contract in
>    site/content/in-dev/unreleased/persistence-consistency.md.
>    - CAS-aware EntityMutation / GrantMutation records.
>    - MetaStoreChangeSet and BasePersistence#commitChangeSet, with backends
>    opting in via supportsAtomicMixedCommit().
>    - commitChangeSet implementations for the in-memory TreeMap backend and
>    JDBC.
>    - AtomicOperationMetaStoreManager and TransactionalMetaStoreManagerImpl
>    grant/revoke paths now build one change set per operation, falling back
> to
>    individual operations when the backend does not support mixed commits.
>    - Tests for the change-set API and TreeMap atomic commits.
>
> I intentionally left out:
>
>    - Entity deletes in MetaStoreChangeSet.
>    - Refactoring createCatalog, dropEntity, and renameEntity to use change
>    sets.
>    - NoSQL support for commitChangeSet.
>
> I am treating this as Phase 1 so the contract and the SPI primitive can be
> reviewed before the larger operation migrations land.
>
> Could you please take a look? In particular, I would like to know whether
> the contract captures the consensus so far, or if it needs to address retry
> / external-work coordination before we merge this foundation.
>
> Thanks,
> Prithvi S
>
> On Mon, Aug 17, 2026 at 7:40 PM Dmitri Bourlatchkov <[email protected]>
> wrote:
>
> > Hi Robert,
> >
> > I agree that the client-visible operation (e.g. REST API request)
> consists
> > of many distinct phases. The database / persistence changes are just one
> of
> > these phases. STS (as an example) is another. My point about transactions
> > in JDBC Persistence applies to the former (database changes) phase - one
> > retry attempt there should ideally involve exactly one transaction.
> >
> > We certainly need logic inside Polaris Servers to to coordinate requests
> to
> > various external systems depending on outcomes from previous request
> > processing phases.
> >
> > I also agree that validation based on database state should be repeated
> > from scratch if we retry the database changes.
> >
> > However, internal RBAC authorization naturally has to happen inside the
> > database update phase (read-check-write). IIRC, that is part of the
> > implicit consistency guarantees, which were discussed during early
> project
> > months (but never got written down in full clarity, unfortunately). If
> the
> > Authorization check is inside that phase for internal RBAC, I suppose it
> > will be there for other Authorizers too and that can involve external
> > systems (e.g. OPA / Ranger). We can certainly expect authorization to be
> > efficient, but I'm not sure how quick it can be in practice. There's
> > certainly potential for delays in some (perhaps infrequent) cases.
> >
> > Cheers,
> > Dmitri.
> >
> >
> > On Mon, Aug 17, 2026 at 9:20 AM Robert Stupp <[email protected]> wrote:
> >
> > > Hi Dmitri,
> > >
> > > I agree that one logical operation needs a consistent view of the whole
> > > backend state.
> > > What worries me is treating one long JDBC transaction as the boundary
> of
> > > that operation and then retrying the complete request if the commit
> > fails.
> > >
> > > The request can also write metadata to object storage or call STS.
> > > Each system can leave us with an outcome whose certainty we cannot
> know.
> > > The database may have committed before the connection was lost, an
> > > object-store write may have succeeded before a timeout, or STS may have
> > > issued credentials before its response was lost.
> > > A database rollback cannot undo any of those effects, and retrying the
> > > whole request may repeat them.
> > >
> > > So I think the contract has to separate the overall operation from each
> > > attempt to commit it.
> > > Each database or backend commit attempt should stay short.
> > > When we retry, validation and authorization need to use the state for
> > that
> > > new attempt, and we need clear rules for which external work can be
> > reused,
> > > repeated, cleaned up, or reconciled after an uncertain outcome.
> > >
> > > Some contextual state could be useful for carrying that state.
> > > But the JDBC connection and transaction should be an implementation
> > detail
> > > inside it, not the boundary of the complete operation.
> > >
> > > Cheers,
> > > Robert
> > >
> > >
> > > On Mon, Aug 10, 2026 at 11:04 PM Dmitri Bourlatchkov <[email protected]
> >
> > > wrote:
> > >
> > > > Hi All,
> > > >
> > > > The idea of a backend-agnostic change-set primitive sounds good.
> > > >
> > > > However, I am not sure it is sufficient for all the use cases we
> > touched
> > > > here. More specifically, the state read by validation code is not
> > > > necessarily reflected in the change set:
> > > >
> > > > * Unchanged entities might be considered by validation, but changed
> in
> > a
> > > > parallel request.
> > > > * Even for changed entities the state read by validation may differ
> > from
> > > > the state being altered. Some validation code talks to the MetaStore
> > > > directly, outside of the data produced by the Resolver.
> > > >
> > > > I do not think Polaris offers any explicit mechanisms (ATM) to ensure
> > > > consistency between these reads and subsequent writes.
> > > >
> > > > As far as JDBC goes, running an overarching Transaction across all
> > reads
> > > > and writes in the same request, at the SERIALIZABLE isolation level
> in
> > > the
> > > > backing RDBMS could solve the problem, I think.
> > > >
> > > > Indeed, such a transaction might be long to accommodate calls made by
> > > > Polaris to external storage, etc. However, is that a problem? I think
> > the
> > > > alternative is for the JDBC Persistence impl. to "manually" track all
> > > reads
> > > > under the same request and redo them in the small transaction that
> > > persists
> > > > the writes. This will also add RDBMS overhead and require complex
> code
> > in
> > > > Polaris to handle the data properly.
> > > >
> > > > Tracking the request-wide transaction can be done only for JDBC
> without
> > > > leaking "transaction" concepts to the NoSQL Persistence, I think.
> NoSQL
> > > > will use other mechanisms to ensure read/write consistency.
> > > >
> > > > Connecting to Robert's email (a parallel branch in this discussion),
> > I'd
> > > > like to propose this approach:
> > > >
> > > > * Each request establishes a "Data Context"
> > > >   - In JDBC the Data Context corresponds to a JDBC Connection + Tx
> > > >   - In NoSQL the Data Context tracks one or more reference hashes
> > > > * All Persistence access in the same request goes through the same
> Data
> > > > Context
> > > > * All Persistence changes are committed once at the end of the
> request
> > > >   - Not all changes have to be globally atomic. For example, changes
> in
> > > > Catalogs A and B do not have to be atomic with respect to each other.
> > We
> > > > can go deeper into this later. This is relevant to NoSQL.
> > > >   - Transactional backends like JDBC can, of course, choose to make
> all
> > > > changes globally atomic.
> > > > * A commit can fail in two main ways:
> > > >   - A retriable failure like an optimistic lock error or Tx
> > > serializability
> > > > error
> > > >   - A non-triable logical error (e.g. entity not found)
> > > > * On a retriable error the whole request is re-attempted (as if
> > > resubmitted
> > > > by a client) a few times (configurable timeout).
> > > >
> > > > WDYT?
> > > >
> > > > Cheers,
> > > > Dmitri.
> > > >
> > > > On Sun, Jul 26, 2026 at 12:44 PM Jean-Baptiste Onofré <
> [email protected]
> > >
> > > > wrote:
> > > >
> > > > > Hi all
> > > > >
> > > > > I think Robert has a good point.
> > > > >
> > > > > If the atomicity guarantee lives only in BasePersistence, then the
> > > > > manager contract can't tell a caller whether the state it read for
> > > > > validation/authorization/credential-vending is the same state that
> > > > > eventually commits. That's the actual bug class behind the JDBC
> > > > > symptoms (not any single operation being non-atomic, but the
> contract
> > > > > being silent about it. Every operation-specific fix (like #4939,
> > #5035
> > > > > or #5095) re-answers this question locally and it keeps recurring.
> So
> > > > > the deliverable should start with a written consistency contract at
> > > > > the manager level.
> > > > >
> > > > > I'm not sure migrating everyone to
> TransactionalMetaStoreManagerImpl
> > > > > is a good idea. It would tie the logical change set to a DB
> > > > > transaction spanning the REST request. It means:
> > > > > - it holds a durable transaction open across slow external work
> > > > > (credential vending, OPA, Ranger, ...)
> > > > > - it doesn't map to NoSQL
> > > > > - It wraps single-row updates in runWiithinTransaction, which is a
> > > > overhead
> > > > >
> > > > > So, I think the transactional manager isn't a portable target. It's
> > > > > "only" one backend's strategy.
> > > > >
> > > > > I think Privthi's approach is right. A backend-agnostic change-set
> > > > > primitive with a documented fallback is the correct shape. There is
> > > > > one caveat: without carrying the original entitiy (not just the new
> > > > > one) the change set can't express optimistic concurrency.
> > > > >
> > > > > I propose the following multi-steps approach:
> > > > > 1. We write the manager-level consistency contract as Robert asked.
> > It
> > > > > should include the explicit statement that a logical change set is
> > not
> > > > > a request-scope DB transaction.
> > > > > 2. We make Compare And Swap (optimistic-concurrency pattern)
> baseline
> > > > > first-class in EntityMutation before merging the SPI
> > > > > 3. Then we refactor createCatalog/dropEntity/renameEntity.
> > > > >
> > > > > Thoughts?
> > > > >
> > > > > Regards
> > > > > JB
> > > > >
> > > > > On Fri, Jul 24, 2026 at 5:08 PM Robert Stupp <[email protected]>
> wrote:
> > > > > >
> > > > > > Hi all,
> > > > > >
> > > > > > Yufei's clarification about where the atomicity guarantee is
> > defined
> > > > > seems
> > > > > > important.
> > > > > > If it is a property of the lower-level BasePersistence contract
> > > rather
> > > > > than
> > > > > > the general PolarisMetaStoreManager contract, the general
> contract
> > > does
> > > > > not
> > > > > > tell callers whether the state used for validation,
> authorization,
> > or
> > > > > > credential vending is consistent with the change that eventually
> > > > commits.
> > > > > >
> > > > > > The current PRs suggest that operation-specific multi-object
> > methods
> > > > can
> > > > > > fix individual cases, while leaving the same contract question to
> > > recur
> > > > > for
> > > > > > each new case.
> > > > > >
> > > > > > I would also avoid defining a logical change set as a database
> > > > > transaction
> > > > > > around an entire REST request.
> > > > > > That would tie the contract to the backend and database, and
> could
> > > keep
> > > > > the
> > > > > > durable attempt open across slow external work.
> > > > > >
> > > > > > So is the choice really between the two current manager
> > > > implementations,
> > > > > or
> > > > > > do we first need to revisit the boundary and guarantees exposed
> to
> > > > their
> > > > > > callers?
> > > > > >
> > > > > > Cheers,
> > > > > > Robert
> > > > > >
> > > > > >
> > > > > > On Fri, Jul 17, 2026 at 10:21 PM Dmitri Bourlatchkov <
> > > [email protected]
> > > > >
> > > > > > wrote:
> > > > > >
> > > > > > > Hi Yufei,
> > > > > > >
> > > > > > > Thanks for the link. I stand corrected. The "one atomic change
> > per
> > > > > method"
> > > > > > > contract as defined in javadoc does apply only to
> BasePersistence
> > > > > > > and AtomicOperationMetaStoreManager (which delegates to
> > > > > BasePersistence).
> > > > > > >
> > > > > > > Note that all non-test Persistence implementations in the
> Polaris
> > > > > codebase
> > > > > > > extend those classes (which is probably why I was confused
> about
> > > > > atomicity
> > > > > > > expectations).
> > > > > > >
> > > > > > > However, this creates a gap in the Persistence SPI
> specification.
> > > If
> > > > > > > other PolarisMetaStoreManager implementations do not have to
> > comply
> > > > > with
> > > > > > > this principle, it will create a conceptual difficulty at call
> > > sites.
> > > > > How
> > > > > > > can PolarisMetaStoreManager callers reason about consistency
> and
> > > > > durability
> > > > > > > behaviours in general?
> > > > > > >
> > > > > > > I believe we need to address that as part of this discussion.
> > > > > > >
> > > > > > > Cheers,
> > > > > > > Dmitri.
> > > > > > >
> > > > > > > On Thu, Jul 16, 2026 at 9:29 PM Yufei Gu <[email protected]
> >
> > > > wrote:
> > > > > > >
> > > > > > > > > The MetaStore SPI is currently defined with the idea that
> one
> > > > > method
> > > > > > > call
> > > > > > > > means one atomic change.
> > > > > > > >
> > > > > > > > If the "MetaStore SPI" refers to the interface
> > > > > PolarisMetaStoreManager, I
> > > > > > > > don't think we've ever state each method to be atomic. We did
> > > > clarify
> > > > > > > > atomicity[1] in the interface BasePersistence though.
> > > > > > > >
> > > > > > > >
> > > > > > > > 1.
> > > > > > > >
> > > > > > > >
> > > > > > >
> > > > >
> > > >
> > >
> >
> https://github.com/apache/polaris/blob/e9039e12003a13e783b5130a3d30d30cfe78d93c/polaris-core/src/main/java/org/apache/polaris/core/persistence/BasePersistence.java#L48
> > > > > > > >
> > > > > > > >
> > > > > > > >
> > > > > > > > Yufei
> > > > > > > >
> > > > > > > >
> > > > > > > > On Thu, Jul 16, 2026 at 11:27 AM Dmitri Bourlatchkov <
> > > > > [email protected]>
> > > > > > > > wrote:
> > > > > > > >
> > > > > > > > > Hi Yufei,
> > > > > > > > >
> > > > > > > > > I agree that JDBC transactions must be handled more
> > explicitly.
> > > > > > > However,
> > > > > > > > > I'm not sure that simply moving to
> > > > > TransactionalMetaStoreManagerImpl is
> > > > > > > > > sufficient.
> > > > > > > > >
> > > > > > > > > The MetaStore SPI is currently defined with the idea that
> one
> > > > > method
> > > > > > > call
> > > > > > > > > means one atomic change [1]. The "transactional" MetaStore
> > > impl.
> > > > is
> > > > > > > but a
> > > > > > > > > sub-case of that. It cannot alter the high-level contract.
> > > > > > > > >
> > > > > > > > > We could add SPI methods having multiple object parameters
> to
> > > > > represent
> > > > > > > > > grouped changes, but I am not sure it will be a sound
> design.
> > > > This
> > > > > will
> > > > > > > > > bloat the interface surfaces and require extra impl. effort
> > for
> > > > > each
> > > > > > > > > backend type. More importantly, adding multi-arg change
> > methods
> > > > > still
> > > > > > > > won't
> > > > > > > > > address the problem of reads being consistent with writes,
> > > > because
> > > > > each
> > > > > > > > > method call will still be independent regarding the data
> > stored
> > > > in
> > > > > the
> > > > > > > > > database.
> > > > > > > > >
> > > > > > > > > I tend to think we need to introduce a "change set" or
> > "atomic
> > > > > batch"
> > > > > > > > > concept to core Persistence and associate each REST API
> > request
> > > > > with
> > > > > > > one
> > > > > > > > > such change set, which will be committed (or rolled back)
> at
> > > the
> > > > > end of
> > > > > > > > the
> > > > > > > > > request. I believe Ayush mentioned a similar concept in PR
> > 4939
> > > > > [2]. In
> > > > > > > > > JDBC each change set will naturally be associated with an
> > RDBMS
> > > > > > > > > transaction. In NoSQL persistence, each atomic change set
> > will
> > > be
> > > > > > > > > associated with one CAS operation on the underlying
> database.
> > > > > > > > >
> > > > > > > > > [1]
> > > > > https://lists.apache.org/thread/rf5orxs815zs4h64p4rwp03q3pbgxb5r
> > > > > > > > >
> > > > > > > > > [2]
> > > > > https://github.com/apache/polaris/pull/4939#discussion_r3575719158
> > > > > > > > >
> > > > > > > > > Cheers,
> > > > > > > > > Dmitri.
> > > > > > > > >
> > > > > > > > > On Thu, Jul 16, 2026 at 1:12 PM Yufei Gu <
> > [email protected]
> > > >
> > > > > wrote:
> > > > > > > > >
> > > > > > > > > > Thanks for raising this, Dmitri.
> > > > > > > > > >
> > > > > > > > > > These are valid concerns, and they were already
> recognized
> > > when
> > > > > we
> > > > > > > > > > introduced JDBC persistence to Polaris. At that time, we
> > > chose
> > > > > to use
> > > > > > > > > > AtomicOperationMetaStoreManager for the JDBC due to the
> > > > > simplicity. I
> > > > > > > > > think
> > > > > > > > > > most of the issues mentioned here can already be
> addressed
> > by
> > > > > > > > > > TransactionalMetaStoreManagerImpl.
> > > > > > > > > >
> > > > > > > > > > For example, rename is already wrapped in a transaction
> in
> > > > > > > > > > TransactionalMetaStoreManagerImpl [1]. Similarly, catalog
> > > > > creation,
> > > > > > > > which
> > > > > > > > > > involves reading and creating multiple objects, is also
> > > > executed
> > > > > > > > within a
> > > > > > > > > > transaction [2].
> > > > > > > > > >
> > > > > > > > > > I see two possible directions:
> > > > > > > > > >
> > > > > > > > > >    1.
> > > > > > > > > >
> > > > > > > > > >    Modify AtomicOperationMetaStoreManager together with
> the
> > > > > > > persistence
> > > > > > > > > >    backends (such as JDBC) to provide the required
> > > consistency
> > > > > > > > guarantees
> > > > > > > > > > for
> > > > > > > > > >    specific operations, similar to what
> > > > > > > > TransactionalMetaStoreManagerImpl
> > > > > > > > > >    does.
> > > > > > > > > >    2.
> > > > > > > > > >
> > > > > > > > > >    Migrate the persistence backends (such as JDBC) to use
> > > > > > > > > >    TransactionalMetaStoreManagerImpl directly. We may
> have
> > to
> > > > > deal
> > > > > > > with
> > > > > > > > > >    transactional semantic mismatches across different
> > > > persistence
> > > > > > > > > backends.
> > > > > > > > > >    For example, we would likely avoid using JDBC's
> > > > > > > > `runWithinTransaction`
> > > > > > > > > > for
> > > > > > > > > >    single row updates, which adds additional overhead and
> > > > > complexity
> > > > > > > > > > without
> > > > > > > > > >    benefits.
> > > > > > > > > >
> > > > > > > > > > References:
> > > > > > > > > >
> > > > > > > > > >    1.
> > > > > > > > > >
> > > > > > > > > >
> > > > > > > > > >
> > > > > > > > > >
> > > > > > > > >
> > > > > > > >
> > > > > > >
> > > > >
> > > >
> > >
> >
> https://github.com/apache/polaris/blob/5731c5cbee02257d1f21f78ca3befcd639b100a3/polaris-core/src/main/java/org/apache/polaris/core/persistence/transactional/TransactionalMetaStoreManagerImpl.java#L1286
> > > > > > > > > >    2.
> > > > > > > > > >
> > > > > > > > > >
> > > > > > > > > >
> > > > > > > > > >
> > > > > > > > >
> > > > > > > >
> > > > > > >
> > > > >
> > > >
> > >
> >
> https://github.com/apache/polaris/blob/5731c5cbee02257d1f21f78ca3befcd639b100a3/polaris-core/src/main/java/org/apache/polaris/core/persistence/transactional/TransactionalMetaStoreManagerImpl.java#L965
> > > > > > > > > >
> > > > > > > > > > Yufei
> > > > > > > > > >
> > > > > > > > > >
> > > > > > > > > > On Thu, Jul 16, 2026 at 8:22 AM Dmitri Bourlatchkov <
> > > > > > > [email protected]>
> > > > > > > > > > wrote:
> > > > > > > > > >
> > > > > > > > > > > Hi all,
> > > > > > > > > > >
> > > > > > > > > > > Ayush and Prithvi recently contributed a couple of
> > > > interesting
> > > > > PRs:
> > > > > > > > > > > [4939], [5035].
> > > > > > > > > > >
> > > > > > > > > > > It looks like people are starting to encounter
> > consistency
> > > > > issues
> > > > > > > in
> > > > > > > > > > > JDBC persistence.
> > > > > > > > > > >
> > > > > > > > > > > The PRs provide valuable insight into the underlying
> > > issues.
> > > > > They
> > > > > > > > offer
> > > > > > > > > > > incremental fixes that can work. However, I believe it
> is
> > > > time
> > > > > for
> > > > > > > > the
> > > > > > > > > > > Polaris community to review and improve this area of
> the
> > > > > codebase
> > > > > > > > > > > holistically.
> > > > > > > > > > >
> > > > > > > > > > > By this, I mean finding a solution that can be applied
> to
> > > all
> > > > > > > > > > > persistence backends (in-memory, JDBC, NoSQL) and
> > addresses
> > > > > these
> > > > > > > > > > > aspects:
> > > > > > > > > > >
> > > > > > > > > > > * Supporting concurrent and consistent changes where
> the
> > > > > service
> > > > > > > > reads
> > > > > > > > > > >   and validates current catalog state, then commits a
> > > change
> > > > > (e.g.
> > > > > > > > > > >   name clashes during renames).
> > > > > > > > > > > * Supporting consistent but independent changes to RBAC
> > > > grants
> > > > > and
> > > > > > > > > > >   MetaStore entities. This independence is needed to
> > > support
> > > > > > > > > > >   external authorizers like OPA and Ranger.
> > > > > > > > > > > * Supporting atomic changes across multiple similar
> > > entities.
> > > > > > > > > > > * Supporting authorization-based filtering of list
> > > operations
> > > > > (cf.
> > > > > > > > > > >   [4831]).
> > > > > > > > > > > * Supporting credential-vending decisions that are
> rooted
> > > in
> > > > > the
> > > > > > > > > > >   exact state of the catalog.
> > > > > > > > > > > * Supporting server-side retries for transient
> > persistence
> > > > > failures
> > > > > > > > > > >   (e.g. RDBMS Tx serializability failures).
> > > > > > > > > > >
> > > > > > > > > > > Please share your comments and ideas.
> > > > > > > > > > >
> > > > > > > > > > > [4831] https://github.com/apache/polaris/pull/4831
> > > > > > > > > > >
> > > > > > > > > > > [4939] https://github.com/apache/polaris/pull/4939
> > > > > > > > > > >
> > > > > > > > > > > [5035] https://github.com/apache/polaris/pull/5035
> > > > > > > > > > >
> > > > > > > > > > > Thanks,
> > > > > > > > > > > Dmitri
> > > > > > > > > > >
> > > > > > > > > >
> > > > > > > > >
> > > > > > > >
> > > > > > >
> > > > >
> > > >
> > >
> >
>

Reply via email to