Hi All,

(replying partially)

I very much support Robert's proposal for using a well-defined and
unambiguous format for namespaces.

Many Polaris APIs fall into following the IRC approach to namespace
representation in query parameters. Yet, that approach has multiple issues
, which can be seen in Iceberg dev ML / GH issues.

I think Polaris should use a more robust namespace representation in its
native APIs.

Cheers,
Dmitri.

On Tue, Sep 8, 2026 at 6:41 AM Robert Stupp <[email protected]> wrote:

> Hi,
>
> I have been thinking about what a full implementation would require.
>
> I like that the proposal separates definitions, assignments, and effective
> reads. I have a few API-contract questions that seem worth resolving while
> the contract is still separate from the implementation.
>
>
> First, I think the target query parameters need a defined encoding for
> identifier elements.
>
> This is not only about the unit separator. Namespace elements and object
> names may themselves contain characters such as &, ?, =, +, or %.
> Those must be preserved rather than interpreted as query syntax.
>
> Ordinary URI query-value encoding handles those characters, but it does not
> solve the separate problem of representing the boundaries between multipart
> namespace elements.
>
> I suggest defining a small, reversible namespace-element codec, then
> applying ordinary URI encoding to its complete output. Nessie's escaped
> path
> representation is a useful precedent: it has an unambiguous element
> separator
> and escape syntax, while avoiding control characters in the transport
> representation.
>
> The contract should specify the codec, its decoding failures, and
> conformance examples. Client libraries should expose it rather than
> requiring
> every client to reproduce it. The structured target used by the write APIs
> would still be the clearest canonical representation; this codec would make
> the GET form safe and interoperable.
>
>
> Second, I think the revision token should be opaque at the API boundary.
>
> The backend should be free to use a native row revision, commit ID, ETag,
> or
> another conditional-write token. However, the contract should define the
> observable precondition: the server returns a token, and an update succeeds
> only if the client supplies the token for the current tag definition.
> Otherwise the server returns a conflict.
>
> That requires token matching semantics, but not an integer type, an initial
> value, ordering, increment-by-one behavior, or history semantics. A
> catalog-wide commit token would also be valid, although it could create
> avoidable conflicts for unrelated changes.
>
>
> Third, the direct reverse lookup is useful, but I would treat it as a
> first-class, paginated relationship rather than a tag record containing a
> collection of targets.
>
> A common tag can legitimately be attached to a very large number of objects
> or columns. A backend will normally need one forward access path for direct
> assignments by target, and one reverse access path by tag/value, with
> backend-specific partitioning or sharding. Effective assignments should
> remain
> computed from the target and its ancestors; materializing inherited
> assignments onto descendants would have very different scaling behavior.
>
> This also affects detach-all. Deleting an unbounded number of assignment
> records atomically is not a portable primitive for all backends. The
> contract
> should distinguish observable deletion semantics from physical cleanup, or
> state the backend capability required for a synchronous detach-all
> operation.
>
>
> Finally, I agree with the V1 boundary of top-level Iceberg columns, but I
> would
> keep the core tag model independent of Iceberg. For Iceberg, the durable
> column reference should be the field ID, with a column name used only for
> request-time resolution and display. Other table implementations could opt
> in
> later once they provide an equally stable field identity. This avoids
> treating
> a name-based column mapping as a general abstraction.
>
>
> None of this requires tags to become authorization inputs in V1. It is
> mainly
> about leaving the assignment and read contract implementable by more than
> one
> persistence model when those slices arrive.
>
> Thanks,
> Robert
>
>
> On Fri, Aug 28, 2026 at 2:19 AM EJ Wang <[email protected]>
> wrote:
>
> > Hi folks,
> >
> > A quick update on the Tag work. *Current status*:
> > - PR1: API contract (https://github.com/apache/polaris/pull/5366): ready
> > for review
> > - *NEW! *PR2: Definition CRUD (
> https://github.com/apache/polaris/pull/5391
> > ):
> > open as Draft, ready for review if PR1 LGTY
> > - PR3: Assignment writes and storage: planned
> > - PR4: Reads, inheritance, and reverse lookup: planned
> >
> > *More on PR2:*
> > - PR2 makes Tag definitions usable through create, list, load, update,
> > rename, and delete. It intentionally stops before assignments, so their
> > persistence model remains open for the next slice.
> > - The PR is stacked on #5366 and will be rebased once that PR merges.
> >
> > *Asks:*
> > - For #5366, please call out any remaining API contract concerns. For
> > #5391, I would especially appreciate feedback on the slice boundary and
> the
> > decision to reuse the existing entity persistence model.
> > - The design doc remains here:
> >
> >
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
> >
> > I’ll keep using this thread for new delivery slices, material status
> > changes, and specific community asks.
> >
> > Thanks,
> > -ej
> >
> > On Mon, Aug 24, 2026 at 5:04 PM EJ Wang <[email protected]>
> > wrote:
> >
> > > Hi folks,
> > >
> > > Following up on this thread, I have opened a PR to land the public API
> > > contract for Tags: https://github.com/apache/polaris/pull/5366
> > >
> > > The PR defines Tag management, assignment and unassignment, direct and
> > > inherited reads, and reverse lookup. V1 covers catalogs, namespaces,
> > > Iceberg and generic tables as whole objects, and top-level Iceberg
> table
> > > columns. Views, generic-table columns, nested fields, multi-value
> > > assignments, and tag-based authorization are deferred.
> > >
> > > I plan to deliver the capability through four PRs that merge in order:
> > the
> > > API contract in this PR, Tag definition CRUD, assignment writes and
> > > storage, then reads and reverse lookup. A separate follow-up will add
> > > grants on Tag resources to the management APIs. That grant surface is
> > > distinct from using Tags to control access to tagged objects, which
> > remains
> > > outside v1.
> > >
> > > The updated design doc is here:
> > >
> >
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
> > >
> > > The PR is currently Draft while we finish aligning on the public
> > contract.
> > > It is intended to merge as the first delivery slice, not remain as a
> > > design-only artifact. Please call out any remaining scope or contract
> > > concerns. If the list is aligned, I will mark it ready for review.
> > >
> > > Thanks,
> > > -ej
> > >
> > > On Wed, Aug 12, 2026 at 2:05 PM EJ Wang <
> [email protected]>
> > > wrote:
> > >
> > >> Thanks Dmitri, these comments were very useful.
> > >>
> > >> I went through the three areas you called out and updated the proposal
> > >> accordingly.
> > >>
> > >> On the permission/policy direction, *I agree the Tag model should
> leave
> > >> room for permissions or policies to consume tags later*, including the
> > >> direction JB proposed. I am keeping that outside the v1 Tag contract,
> > >> though. In v1, tags classify resources; they do not themselves grant
> or
> > >> deny access. Polaris Policy looks like the closest existing foundation
> > if
> > >> we later want a portable tag-aware policy model, but I think that
> > deserves
> > >> a separate proposal rather than baking policy semantics into the Tag
> > >> storage model now.
> > >>
> > >> I also made the authorizer path more explicit. *A future OPA, Ranger,
> or
> > >> other authorizer could receive the target's complete effective tags as
> > >> resource attributes*. The authorization path would resolve those tags
> > >> internally, applying target-types, inheritance, closest-wins,
> > grandfathered
> > >> values, and the same coherent-read guarantees as the Tag API. At
> > minimum,
> > >> the portable input can include the tag definition ID, current name,
> and
> > >> selected value; provenance can be additional context. If Polaris
> cannot
> > >> resolve the complete effective state, authorization should fail closed
> > >> rather than treat the resource as untagged.
> > >>
> > >> That also makes the persistence expectation on the read path clearer:
> an
> > >> implementation needs to resolve the target and relevant ancestors,
> > obtain
> > >> the applicable tag definitions and assignments, and produce one
> coherent
> > >> effective result. *Those observable semantics are the backend
> contract;
> > >> the physical lookup/indexing strategy is not.*
> > >>
> > >> On the Java interface suggestion, I added Java-shaped records for the
> > >> durable logical model so the definition, target identity, and
> assignment
> > >> shapes are easier to review from JDBC and NoSQL perspectives. I
> stopped
> > >> short of proposing operation interfaces in pseudo-code, though. My
> > current
> > >> thinking is that we should first agree on the durable facts and
> required
> > >> behavior, then design the actual persistence SPI around the needs of
> the
> > >> implementations. I did not want an illustrative interface in this
> > design to
> > >> accidentally become the persistence contract.
> > >>
> > >> So Part 2 now separates the two intentionally:
> > >>
> > >> *logical data + behavior/conformance requirements are specified;
> > >> transaction, CAS, atomic batch, provider-native operations, and the
> > >> eventual Java SPI remain implementation/design choices.*
> > >>
> > >> Thanks again for the review, and definitely keep the comments coming
> :)
> > >>
> > >> I've updated the doc, please check it out the latest and the greatest:
> > >>
> > >>
> >
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0
> > >>
> > >> -ej
> > >>
> > >> On Fri, Aug 7, 2026 at 3:39 PM Dmitri Bourlatchkov <[email protected]>
> > >> wrote:
> > >>
> > >>> Hi EJ, JB,
> > >>>
> > >>> I left some comments on EJ's doc.  I actually have a lot of comments
> on
> > >>> the
> > >>> REST API design, I only posted some of them to start a discussion
> > >>> without overloading the doc.
> > >>>
> > >>> Overall, I believe EJ's proposal should also allow permission
> > assignments
> > >>> on tags that JB proposed (eventually). We just need to clearly define
> > the
> > >>> persistence expectations for looking up related tags on the read
> path.
> > >>>
> > >>> We should probably specify whether and how tags are exposed to
> > >>> authorizers
> > >>> (OPA, Ranger). I imagine people will want to use them in external
> > policy
> > >>> engines the moment the feature is available.
> > >>>
> > >>> On the persistence side, I believe it would be nice to define actual
> > java
> > >>> interfaces (perhaps in pseudo code) to allow easier review from the
> > NoSQL
> > >>> persistence perspective (also commented in the doc).
> > >>>
> > >>> Cheers,
> > >>> Dmitri.
> > >>>
> > >>> On Thu, Jul 30, 2026 at 12:39 AM Jean-Baptiste Onofré <
> [email protected]
> > >
> > >>> wrote:
> > >>>
> > >>> > Hi EJ
> > >>> >
> > >>> > Thanks for starting this discussion.
> > >>> >
> > >>> > For the record, here's my initial proposal about tagging:
> > >>> > https://lists.apache.org/thread/nmqmmjfmocfllb71fcmyp9syc9gyn820
> > >>> >
> > >>> > At that time, only Dmitri replied :)
> > >>> > So, I would be happy to work with you on this, as I still have the
> > PoC
> > >>> > I created for my initial proposal.
> > >>> >
> > >>> > I will try to join the scheduled meeting (no guarantee).
> > >>> >
> > >>> > Regards
> > >>> > JB
> > >>> >
> > >>> > On Fri, Jul 17, 2026 at 6:53 AM EJ Wang <
> > >>> [email protected]>
> > >>> > wrote:
> > >>> > >
> > >>> > > Hi folks,
> > >>> > >
> > >>> > > I have prepared a Google Doc
> > >>> > > <
> > >>> >
> > >>>
> >
> https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing
> > >>> > >
> > >>> > > for the Polaris tag spec proposal.
> > >>> > >
> > >>> > > The goal is simple: add a native tag model to Polaris so users
> can
> > >>> > classify
> > >>> > > catalog objects, read those classifications back, and find
> objects
> > by
> > >>> > tag.
> > >>> > >
> > >>> > > The proposal covers:
> > >>> > > * tag definitions as catalog-scoped Polaris entities
> > >>> > > * tag assignments on catalogs, namespaces, table-like objects,
> and
> > >>> > columns
> > >>> > > * allowed values on tag definitions
> > >>> > > * direct and inherited tag reads
> > >>> > > * direct by-tag lookup
> > >>> > > * the durable model behind the API
> > >>> > > * how this compares with the existing Polaris Policy API (tag
> > design
> > >>> > > referenced policy heavily, given their pattern similarity)
> > >>> > >
> > >>> > > Please take a look and leave comments in the doc. Let me know
> WDYT!
> > >>> > >
> > >>> > > I would also like to discuss this in the July 23 community sync.
> A
> > >>> > separate
> > >>> > > dedicated review meeting will be scheduled separately, likely
> > within
> > >>> the
> > >>> > > next two weeks.
> > >>> > >
> > >>> > > Thanks,
> > >>> > > -ej
> > >>> >
> > >>>
> > >>
> >
>

Reply via email to