Hi All,

I posted a comment about tags API URI paths that may be of interest to a
few people. Re-posting here for awareness:

https://github.com/apache/polaris/pull/5366#discussion_r3980198150

Cheers,
Dmitri.

On Thu, Sep 10, 2026 at 1:21 AM EJ Wang <[email protected]>
wrote:

> Hi folks,
>
> A quick update on the Tag work. PR1-3 and the design doc have been updated
> following the discussion.
>
> *Current status*:
> PR1: API contract: ready for review.
> PR2: Definition CRUD: open as Draft, ready for review if PR1 LGTY.
> *NEW! *PR3: Assignment writes and storage: open as Draft, ready for review
> if PR2 LGTY.
> PR4: Reads, inheritance, and reverse lookup: planned.
>
> *More on PR3*:
> PR3 adds assign/unassign for catalogs, namespaces, tables, and top-level
> Iceberg columns, plus atomic detach-all. It includes JDBC and in-memory
> support. Reads remain in PR4.
> The PR is stacked on PR2. Its description links the assignment-only commit
> for focused review.
>
> *Asks*:
> For PR3, I’d especially appreciate feedback on the persistence interfaces,
> their impact on external implementations, and the detach-all guarantee.
> The updated design doc also includes the reverse-lookup flow and a
> discussion of the deferred shared encoding work.
>
> Thanks,
> -ej
>
> On Wed, Sep 9, 2026 at 3:09 PM EJ Wang <[email protected]>
> wrote:
>
> > Thanks Robert, Prithvi, and Dmitri. I’ve updated the spec doc
> > <
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0#heading=h.jx650zq2bd2g>
> with
> > the contract clarifications (see below) and the follow-up discussion with
> > Dmitri. The corresponding PR updates are underway.
> >
> > *For encoding*, my proposal is to retain Iceberg’s namespace query
> > convention in v1, with explicit supported-name, encoding, and decoding
> > rules in Part 1 section 5.2. This retains the known separator and ingress
> > limitations. Since the issue also affects existing Iceberg/Polaris APIs,
> > I’d address a replacement codec in a separate issue/PR. Part 3 section
> 7.12
> > records alternatives and spike results to start that discussion.
> >
> > *Version tokens* are opaque strings in both responses and update
> > requests. Clients return the token unchanged, and stale updates return
> 409.
> > The check must cover every supported definition-write path, while
> backends
> > choose their revision mechanism.
> >
> > *For detach-all*, the definition and assignments must disappear together
> > through Tag reads, or nothing changes. Physical cleanup may follow. An
> > implementation unable to provide that guarantee returns 501 after
> > authorization and before changing visible state.
> >
> > *Column assignments* use table identity and column identifier, using
> > Iceberg field id for Iceberg tables. Renames preserve assignments when
> > identity is preserved, while same-name replacements do not inherit them.
> V1
> > remains limited to top-level Iceberg columns.
> >
> > Does this scope split work for you, particularly keeping the shared codec
> > redesign separate from Tag v1? I’d like to settle these contract points
> > here before PR1 merges.
> >
> > -ej
> >
> > On Wed, Sep 9, 2026 at 7:36 AM Dmitri Bourlatchkov <[email protected]>
> > wrote:
> >
> >> Hi All,
> >>
> >> (replying partially)
> >>
> >> I very much support Robert's proposal for using a well-defined and
> >> unambiguous format for namespaces.
> >>
> >> Many Polaris APIs fall into following the IRC approach to namespace
> >> representation in query parameters. Yet, that approach has multiple
> issues
> >> , which can be seen in Iceberg dev ML / GH issues.
> >>
> >> I think Polaris should use a more robust namespace representation in its
> >> native APIs.
> >>
> >> Cheers,
> >> Dmitri.
> >>
> >> On Tue, Sep 8, 2026 at 6:41 AM Robert Stupp <[email protected]> wrote:
> >>
> >> > Hi,
> >> >
> >> > I have been thinking about what a full implementation would require.
> >> >
> >> > I like that the proposal separates definitions, assignments, and
> >> effective
> >> > reads. I have a few API-contract questions that seem worth resolving
> >> while
> >> > the contract is still separate from the implementation.
> >> >
> >> >
> >> > First, I think the target query parameters need a defined encoding for
> >> > identifier elements.
> >> >
> >> > This is not only about the unit separator. Namespace elements and
> object
> >> > names may themselves contain characters such as &, ?, =, +, or %.
> >> > Those must be preserved rather than interpreted as query syntax.
> >> >
> >> > Ordinary URI query-value encoding handles those characters, but it
> does
> >> not
> >> > solve the separate problem of representing the boundaries between
> >> multipart
> >> > namespace elements.
> >> >
> >> > I suggest defining a small, reversible namespace-element codec, then
> >> > applying ordinary URI encoding to its complete output. Nessie's
> escaped
> >> > path
> >> > representation is a useful precedent: it has an unambiguous element
> >> > separator
> >> > and escape syntax, while avoiding control characters in the transport
> >> > representation.
> >> >
> >> > The contract should specify the codec, its decoding failures, and
> >> > conformance examples. Client libraries should expose it rather than
> >> > requiring
> >> > every client to reproduce it. The structured target used by the write
> >> APIs
> >> > would still be the clearest canonical representation; this codec would
> >> make
> >> > the GET form safe and interoperable.
> >> >
> >> >
> >> > Second, I think the revision token should be opaque at the API
> boundary.
> >> >
> >> > The backend should be free to use a native row revision, commit ID,
> >> ETag,
> >> > or
> >> > another conditional-write token. However, the contract should define
> the
> >> > observable precondition: the server returns a token, and an update
> >> succeeds
> >> > only if the client supplies the token for the current tag definition.
> >> > Otherwise the server returns a conflict.
> >> >
> >> > That requires token matching semantics, but not an integer type, an
> >> initial
> >> > value, ordering, increment-by-one behavior, or history semantics. A
> >> > catalog-wide commit token would also be valid, although it could
> create
> >> > avoidable conflicts for unrelated changes.
> >> >
> >> >
> >> > Third, the direct reverse lookup is useful, but I would treat it as a
> >> > first-class, paginated relationship rather than a tag record
> containing
> >> a
> >> > collection of targets.
> >> >
> >> > A common tag can legitimately be attached to a very large number of
> >> objects
> >> > or columns. A backend will normally need one forward access path for
> >> direct
> >> > assignments by target, and one reverse access path by tag/value, with
> >> > backend-specific partitioning or sharding. Effective assignments
> should
> >> > remain
> >> > computed from the target and its ancestors; materializing inherited
> >> > assignments onto descendants would have very different scaling
> behavior.
> >> >
> >> > This also affects detach-all. Deleting an unbounded number of
> assignment
> >> > records atomically is not a portable primitive for all backends. The
> >> > contract
> >> > should distinguish observable deletion semantics from physical
> cleanup,
> >> or
> >> > state the backend capability required for a synchronous detach-all
> >> > operation.
> >> >
> >> >
> >> > Finally, I agree with the V1 boundary of top-level Iceberg columns,
> but
> >> I
> >> > would
> >> > keep the core tag model independent of Iceberg. For Iceberg, the
> durable
> >> > column reference should be the field ID, with a column name used only
> >> for
> >> > request-time resolution and display. Other table implementations could
> >> opt
> >> > in
> >> > later once they provide an equally stable field identity. This avoids
> >> > treating
> >> > a name-based column mapping as a general abstraction.
> >> >
> >> >
> >> > None of this requires tags to become authorization inputs in V1. It is
> >> > mainly
> >> > about leaving the assignment and read contract implementable by more
> >> than
> >> > one
> >> > persistence model when those slices arrive.
> >> >
> >> > Thanks,
> >> > Robert
> >> >
> >> >
> >> > On Fri, Aug 28, 2026 at 2:19 AM EJ Wang <
> [email protected]
> >> >
> >> > wrote:
> >> >
> >> > > Hi folks,
> >> > >
> >> > > A quick update on the Tag work. *Current status*:
> >> > > - PR1: API contract (https://github.com/apache/polaris/pull/5366):
> >> ready
> >> > > for review
> >> > > - *NEW! *PR2: Definition CRUD (
> >> > https://github.com/apache/polaris/pull/5391
> >> > > ):
> >> > > open as Draft, ready for review if PR1 LGTY
> >> > > - PR3: Assignment writes and storage: planned
> >> > > - PR4: Reads, inheritance, and reverse lookup: planned
> >> > >
> >> > > *More on PR2:*
> >> > > - PR2 makes Tag definitions usable through create, list, load,
> update,
> >> > > rename, and delete. It intentionally stops before assignments, so
> >> their
> >> > > persistence model remains open for the next slice.
> >> > > - The PR is stacked on #5366 and will be rebased once that PR
> merges.
> >> > >
> >> > > *Asks:*
> >> > > - For #5366, please call out any remaining API contract concerns.
> For
> >> > > #5391, I would especially appreciate feedback on the slice boundary
> >> and
> >> > the
> >> > > decision to reuse the existing entity persistence model.
> >> > > - The design doc remains here:
> >> > >
> >> > >
> >> >
> >>
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
> >> > >
> >> > > I’ll keep using this thread for new delivery slices, material status
> >> > > changes, and specific community asks.
> >> > >
> >> > > Thanks,
> >> > > -ej
> >> > >
> >> > > On Mon, Aug 24, 2026 at 5:04 PM EJ Wang <
> >> [email protected]>
> >> > > wrote:
> >> > >
> >> > > > Hi folks,
> >> > > >
> >> > > > Following up on this thread, I have opened a PR to land the public
> >> API
> >> > > > contract for Tags: https://github.com/apache/polaris/pull/5366
> >> > > >
> >> > > > The PR defines Tag management, assignment and unassignment, direct
> >> and
> >> > > > inherited reads, and reverse lookup. V1 covers catalogs,
> namespaces,
> >> > > > Iceberg and generic tables as whole objects, and top-level Iceberg
> >> > table
> >> > > > columns. Views, generic-table columns, nested fields, multi-value
> >> > > > assignments, and tag-based authorization are deferred.
> >> > > >
> >> > > > I plan to deliver the capability through four PRs that merge in
> >> order:
> >> > > the
> >> > > > API contract in this PR, Tag definition CRUD, assignment writes
> and
> >> > > > storage, then reads and reverse lookup. A separate follow-up will
> >> add
> >> > > > grants on Tag resources to the management APIs. That grant surface
> >> is
> >> > > > distinct from using Tags to control access to tagged objects,
> which
> >> > > remains
> >> > > > outside v1.
> >> > > >
> >> > > > The updated design doc is here:
> >> > > >
> >> > >
> >> >
> >>
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
> >> > > >
> >> > > > The PR is currently Draft while we finish aligning on the public
> >> > > contract.
> >> > > > It is intended to merge as the first delivery slice, not remain
> as a
> >> > > > design-only artifact. Please call out any remaining scope or
> >> contract
> >> > > > concerns. If the list is aligned, I will mark it ready for review.
> >> > > >
> >> > > > Thanks,
> >> > > > -ej
> >> > > >
> >> > > > On Wed, Aug 12, 2026 at 2:05 PM EJ Wang <
> >> > [email protected]>
> >> > > > wrote:
> >> > > >
> >> > > >> Thanks Dmitri, these comments were very useful.
> >> > > >>
> >> > > >> I went through the three areas you called out and updated the
> >> proposal
> >> > > >> accordingly.
> >> > > >>
> >> > > >> On the permission/policy direction, *I agree the Tag model should
> >> > leave
> >> > > >> room for permissions or policies to consume tags later*,
> including
> >> the
> >> > > >> direction JB proposed. I am keeping that outside the v1 Tag
> >> contract,
> >> > > >> though. In v1, tags classify resources; they do not themselves
> >> grant
> >> > or
> >> > > >> deny access. Polaris Policy looks like the closest existing
> >> foundation
> >> > > if
> >> > > >> we later want a portable tag-aware policy model, but I think that
> >> > > deserves
> >> > > >> a separate proposal rather than baking policy semantics into the
> >> Tag
> >> > > >> storage model now.
> >> > > >>
> >> > > >> I also made the authorizer path more explicit. *A future OPA,
> >> Ranger,
> >> > or
> >> > > >> other authorizer could receive the target's complete effective
> >> tags as
> >> > > >> resource attributes*. The authorization path would resolve those
> >> tags
> >> > > >> internally, applying target-types, inheritance, closest-wins,
> >> > > grandfathered
> >> > > >> values, and the same coherent-read guarantees as the Tag API. At
> >> > > minimum,
> >> > > >> the portable input can include the tag definition ID, current
> name,
> >> > and
> >> > > >> selected value; provenance can be additional context. If Polaris
> >> > cannot
> >> > > >> resolve the complete effective state, authorization should fail
> >> closed
> >> > > >> rather than treat the resource as untagged.
> >> > > >>
> >> > > >> That also makes the persistence expectation on the read path
> >> clearer:
> >> > an
> >> > > >> implementation needs to resolve the target and relevant
> ancestors,
> >> > > obtain
> >> > > >> the applicable tag definitions and assignments, and produce one
> >> > coherent
> >> > > >> effective result. *Those observable semantics are the backend
> >> > contract;
> >> > > >> the physical lookup/indexing strategy is not.*
> >> > > >>
> >> > > >> On the Java interface suggestion, I added Java-shaped records for
> >> the
> >> > > >> durable logical model so the definition, target identity, and
> >> > assignment
> >> > > >> shapes are easier to review from JDBC and NoSQL perspectives. I
> >> > stopped
> >> > > >> short of proposing operation interfaces in pseudo-code, though.
> My
> >> > > current
> >> > > >> thinking is that we should first agree on the durable facts and
> >> > required
> >> > > >> behavior, then design the actual persistence SPI around the needs
> >> of
> >> > the
> >> > > >> implementations. I did not want an illustrative interface in this
> >> > > design to
> >> > > >> accidentally become the persistence contract.
> >> > > >>
> >> > > >> So Part 2 now separates the two intentionally:
> >> > > >>
> >> > > >> *logical data + behavior/conformance requirements are specified;
> >> > > >> transaction, CAS, atomic batch, provider-native operations, and
> the
> >> > > >> eventual Java SPI remain implementation/design choices.*
> >> > > >>
> >> > > >> Thanks again for the review, and definitely keep the comments
> >> coming
> >> > :)
> >> > > >>
> >> > > >> I've updated the doc, please check it out the latest and the
> >> greatest:
> >> > > >>
> >> > > >>
> >> > >
> >> >
> >>
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0
> >> > > >>
> >> > > >> -ej
> >> > > >>
> >> > > >> On Fri, Aug 7, 2026 at 3:39 PM Dmitri Bourlatchkov <
> >> [email protected]>
> >> > > >> wrote:
> >> > > >>
> >> > > >>> Hi EJ, JB,
> >> > > >>>
> >> > > >>> I left some comments on EJ's doc.  I actually have a lot of
> >> comments
> >> > on
> >> > > >>> the
> >> > > >>> REST API design, I only posted some of them to start a
> discussion
> >> > > >>> without overloading the doc.
> >> > > >>>
> >> > > >>> Overall, I believe EJ's proposal should also allow permission
> >> > > assignments
> >> > > >>> on tags that JB proposed (eventually). We just need to clearly
> >> define
> >> > > the
> >> > > >>> persistence expectations for looking up related tags on the read
> >> > path.
> >> > > >>>
> >> > > >>> We should probably specify whether and how tags are exposed to
> >> > > >>> authorizers
> >> > > >>> (OPA, Ranger). I imagine people will want to use them in
> external
> >> > > policy
> >> > > >>> engines the moment the feature is available.
> >> > > >>>
> >> > > >>> On the persistence side, I believe it would be nice to define
> >> actual
> >> > > java
> >> > > >>> interfaces (perhaps in pseudo code) to allow easier review from
> >> the
> >> > > NoSQL
> >> > > >>> persistence perspective (also commented in the doc).
> >> > > >>>
> >> > > >>> Cheers,
> >> > > >>> Dmitri.
> >> > > >>>
> >> > > >>> On Thu, Jul 30, 2026 at 12:39 AM Jean-Baptiste Onofré <
> >> > [email protected]
> >> > > >
> >> > > >>> wrote:
> >> > > >>>
> >> > > >>> > Hi EJ
> >> > > >>> >
> >> > > >>> > Thanks for starting this discussion.
> >> > > >>> >
> >> > > >>> > For the record, here's my initial proposal about tagging:
> >> > > >>> >
> >> https://lists.apache.org/thread/nmqmmjfmocfllb71fcmyp9syc9gyn820
> >> > > >>> >
> >> > > >>> > At that time, only Dmitri replied :)
> >> > > >>> > So, I would be happy to work with you on this, as I still have
> >> the
> >> > > PoC
> >> > > >>> > I created for my initial proposal.
> >> > > >>> >
> >> > > >>> > I will try to join the scheduled meeting (no guarantee).
> >> > > >>> >
> >> > > >>> > Regards
> >> > > >>> > JB
> >> > > >>> >
> >> > > >>> > On Fri, Jul 17, 2026 at 6:53 AM EJ Wang <
> >> > > >>> [email protected]>
> >> > > >>> > wrote:
> >> > > >>> > >
> >> > > >>> > > Hi folks,
> >> > > >>> > >
> >> > > >>> > > I have prepared a Google Doc
> >> > > >>> > > <
> >> > > >>> >
> >> > > >>>
> >> > >
> >> >
> >>
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
> >> > > >>> > >
> >> > > >>> > > for the Polaris tag spec proposal.
> >> > > >>> > >
> >> > > >>> > > The goal is simple: add a native tag model to Polaris so
> users
> >> > can
> >> > > >>> > classify
> >> > > >>> > > catalog objects, read those classifications back, and find
> >> > objects
> >> > > by
> >> > > >>> > tag.
> >> > > >>> > >
> >> > > >>> > > The proposal covers:
> >> > > >>> > > * tag definitions as catalog-scoped Polaris entities
> >> > > >>> > > * tag assignments on catalogs, namespaces, table-like
> objects,
> >> > and
> >> > > >>> > columns
> >> > > >>> > > * allowed values on tag definitions
> >> > > >>> > > * direct and inherited tag reads
> >> > > >>> > > * direct by-tag lookup
> >> > > >>> > > * the durable model behind the API
> >> > > >>> > > * how this compares with the existing Polaris Policy API
> (tag
> >> > > design
> >> > > >>> > > referenced policy heavily, given their pattern similarity)
> >> > > >>> > >
> >> > > >>> > > Please take a look and leave comments in the doc. Let me
> know
> >> > WDYT!
> >> > > >>> > >
> >> > > >>> > > I would also like to discuss this in the July 23 community
> >> sync.
> >> > A
> >> > > >>> > separate
> >> > > >>> > > dedicated review meeting will be scheduled separately,
> likely
> >> > > within
> >> > > >>> the
> >> > > >>> > > next two weeks.
> >> > > >>> > >
> >> > > >>> > > Thanks,
> >> > > >>> > > -ej
> >> > > >>> >
> >> > > >>>
> >> > > >>
> >> > >
> >> >
> >>
> >
>

Reply via email to