Hi All, I posted a comment about tags API URI paths that may be of interest to a few people. Re-posting here for awareness:
https://github.com/apache/polaris/pull/5366#discussion_r3980198150 Cheers, Dmitri. On Thu, Sep 10, 2026 at 1:21 AM EJ Wang <[email protected]> wrote: > Hi folks, > > A quick update on the Tag work. PR1-3 and the design doc have been updated > following the discussion. > > *Current status*: > PR1: API contract: ready for review. > PR2: Definition CRUD: open as Draft, ready for review if PR1 LGTY. > *NEW! *PR3: Assignment writes and storage: open as Draft, ready for review > if PR2 LGTY. > PR4: Reads, inheritance, and reverse lookup: planned. > > *More on PR3*: > PR3 adds assign/unassign for catalogs, namespaces, tables, and top-level > Iceberg columns, plus atomic detach-all. It includes JDBC and in-memory > support. Reads remain in PR4. > The PR is stacked on PR2. Its description links the assignment-only commit > for focused review. > > *Asks*: > For PR3, I’d especially appreciate feedback on the persistence interfaces, > their impact on external implementations, and the detach-all guarantee. > The updated design doc also includes the reverse-lookup flow and a > discussion of the deferred shared encoding work. > > Thanks, > -ej > > On Wed, Sep 9, 2026 at 3:09 PM EJ Wang <[email protected]> > wrote: > > > Thanks Robert, Prithvi, and Dmitri. I’ve updated the spec doc > > < > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0#heading=h.jx650zq2bd2g> > with > > the contract clarifications (see below) and the follow-up discussion with > > Dmitri. The corresponding PR updates are underway. > > > > *For encoding*, my proposal is to retain Iceberg’s namespace query > > convention in v1, with explicit supported-name, encoding, and decoding > > rules in Part 1 section 5.2. This retains the known separator and ingress > > limitations. Since the issue also affects existing Iceberg/Polaris APIs, > > I’d address a replacement codec in a separate issue/PR. Part 3 section > 7.12 > > records alternatives and spike results to start that discussion. > > > > *Version tokens* are opaque strings in both responses and update > > requests. Clients return the token unchanged, and stale updates return > 409. > > The check must cover every supported definition-write path, while > backends > > choose their revision mechanism. > > > > *For detach-all*, the definition and assignments must disappear together > > through Tag reads, or nothing changes. Physical cleanup may follow. An > > implementation unable to provide that guarantee returns 501 after > > authorization and before changing visible state. > > > > *Column assignments* use table identity and column identifier, using > > Iceberg field id for Iceberg tables. Renames preserve assignments when > > identity is preserved, while same-name replacements do not inherit them. > V1 > > remains limited to top-level Iceberg columns. > > > > Does this scope split work for you, particularly keeping the shared codec > > redesign separate from Tag v1? I’d like to settle these contract points > > here before PR1 merges. > > > > -ej > > > > On Wed, Sep 9, 2026 at 7:36 AM Dmitri Bourlatchkov <[email protected]> > > wrote: > > > >> Hi All, > >> > >> (replying partially) > >> > >> I very much support Robert's proposal for using a well-defined and > >> unambiguous format for namespaces. > >> > >> Many Polaris APIs fall into following the IRC approach to namespace > >> representation in query parameters. Yet, that approach has multiple > issues > >> , which can be seen in Iceberg dev ML / GH issues. > >> > >> I think Polaris should use a more robust namespace representation in its > >> native APIs. > >> > >> Cheers, > >> Dmitri. > >> > >> On Tue, Sep 8, 2026 at 6:41 AM Robert Stupp <[email protected]> wrote: > >> > >> > Hi, > >> > > >> > I have been thinking about what a full implementation would require. > >> > > >> > I like that the proposal separates definitions, assignments, and > >> effective > >> > reads. I have a few API-contract questions that seem worth resolving > >> while > >> > the contract is still separate from the implementation. > >> > > >> > > >> > First, I think the target query parameters need a defined encoding for > >> > identifier elements. > >> > > >> > This is not only about the unit separator. Namespace elements and > object > >> > names may themselves contain characters such as &, ?, =, +, or %. > >> > Those must be preserved rather than interpreted as query syntax. > >> > > >> > Ordinary URI query-value encoding handles those characters, but it > does > >> not > >> > solve the separate problem of representing the boundaries between > >> multipart > >> > namespace elements. > >> > > >> > I suggest defining a small, reversible namespace-element codec, then > >> > applying ordinary URI encoding to its complete output. Nessie's > escaped > >> > path > >> > representation is a useful precedent: it has an unambiguous element > >> > separator > >> > and escape syntax, while avoiding control characters in the transport > >> > representation. > >> > > >> > The contract should specify the codec, its decoding failures, and > >> > conformance examples. Client libraries should expose it rather than > >> > requiring > >> > every client to reproduce it. The structured target used by the write > >> APIs > >> > would still be the clearest canonical representation; this codec would > >> make > >> > the GET form safe and interoperable. > >> > > >> > > >> > Second, I think the revision token should be opaque at the API > boundary. > >> > > >> > The backend should be free to use a native row revision, commit ID, > >> ETag, > >> > or > >> > another conditional-write token. However, the contract should define > the > >> > observable precondition: the server returns a token, and an update > >> succeeds > >> > only if the client supplies the token for the current tag definition. > >> > Otherwise the server returns a conflict. > >> > > >> > That requires token matching semantics, but not an integer type, an > >> initial > >> > value, ordering, increment-by-one behavior, or history semantics. A > >> > catalog-wide commit token would also be valid, although it could > create > >> > avoidable conflicts for unrelated changes. > >> > > >> > > >> > Third, the direct reverse lookup is useful, but I would treat it as a > >> > first-class, paginated relationship rather than a tag record > containing > >> a > >> > collection of targets. > >> > > >> > A common tag can legitimately be attached to a very large number of > >> objects > >> > or columns. A backend will normally need one forward access path for > >> direct > >> > assignments by target, and one reverse access path by tag/value, with > >> > backend-specific partitioning or sharding. Effective assignments > should > >> > remain > >> > computed from the target and its ancestors; materializing inherited > >> > assignments onto descendants would have very different scaling > behavior. > >> > > >> > This also affects detach-all. Deleting an unbounded number of > assignment > >> > records atomically is not a portable primitive for all backends. The > >> > contract > >> > should distinguish observable deletion semantics from physical > cleanup, > >> or > >> > state the backend capability required for a synchronous detach-all > >> > operation. > >> > > >> > > >> > Finally, I agree with the V1 boundary of top-level Iceberg columns, > but > >> I > >> > would > >> > keep the core tag model independent of Iceberg. For Iceberg, the > durable > >> > column reference should be the field ID, with a column name used only > >> for > >> > request-time resolution and display. Other table implementations could > >> opt > >> > in > >> > later once they provide an equally stable field identity. This avoids > >> > treating > >> > a name-based column mapping as a general abstraction. > >> > > >> > > >> > None of this requires tags to become authorization inputs in V1. It is > >> > mainly > >> > about leaving the assignment and read contract implementable by more > >> than > >> > one > >> > persistence model when those slices arrive. > >> > > >> > Thanks, > >> > Robert > >> > > >> > > >> > On Fri, Aug 28, 2026 at 2:19 AM EJ Wang < > [email protected] > >> > > >> > wrote: > >> > > >> > > Hi folks, > >> > > > >> > > A quick update on the Tag work. *Current status*: > >> > > - PR1: API contract (https://github.com/apache/polaris/pull/5366): > >> ready > >> > > for review > >> > > - *NEW! *PR2: Definition CRUD ( > >> > https://github.com/apache/polaris/pull/5391 > >> > > ): > >> > > open as Draft, ready for review if PR1 LGTY > >> > > - PR3: Assignment writes and storage: planned > >> > > - PR4: Reads, inheritance, and reverse lookup: planned > >> > > > >> > > *More on PR2:* > >> > > - PR2 makes Tag definitions usable through create, list, load, > update, > >> > > rename, and delete. It intentionally stops before assignments, so > >> their > >> > > persistence model remains open for the next slice. > >> > > - The PR is stacked on #5366 and will be rebased once that PR > merges. > >> > > > >> > > *Asks:* > >> > > - For #5366, please call out any remaining API contract concerns. > For > >> > > #5391, I would especially appreciate feedback on the slice boundary > >> and > >> > the > >> > > decision to reuse the existing entity persistence model. > >> > > - The design doc remains here: > >> > > > >> > > > >> > > >> > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing > >> > > > >> > > I’ll keep using this thread for new delivery slices, material status > >> > > changes, and specific community asks. > >> > > > >> > > Thanks, > >> > > -ej > >> > > > >> > > On Mon, Aug 24, 2026 at 5:04 PM EJ Wang < > >> [email protected]> > >> > > wrote: > >> > > > >> > > > Hi folks, > >> > > > > >> > > > Following up on this thread, I have opened a PR to land the public > >> API > >> > > > contract for Tags: https://github.com/apache/polaris/pull/5366 > >> > > > > >> > > > The PR defines Tag management, assignment and unassignment, direct > >> and > >> > > > inherited reads, and reverse lookup. V1 covers catalogs, > namespaces, > >> > > > Iceberg and generic tables as whole objects, and top-level Iceberg > >> > table > >> > > > columns. Views, generic-table columns, nested fields, multi-value > >> > > > assignments, and tag-based authorization are deferred. > >> > > > > >> > > > I plan to deliver the capability through four PRs that merge in > >> order: > >> > > the > >> > > > API contract in this PR, Tag definition CRUD, assignment writes > and > >> > > > storage, then reads and reverse lookup. A separate follow-up will > >> add > >> > > > grants on Tag resources to the management APIs. That grant surface > >> is > >> > > > distinct from using Tags to control access to tagged objects, > which > >> > > remains > >> > > > outside v1. > >> > > > > >> > > > The updated design doc is here: > >> > > > > >> > > > >> > > >> > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing > >> > > > > >> > > > The PR is currently Draft while we finish aligning on the public > >> > > contract. > >> > > > It is intended to merge as the first delivery slice, not remain > as a > >> > > > design-only artifact. Please call out any remaining scope or > >> contract > >> > > > concerns. If the list is aligned, I will mark it ready for review. > >> > > > > >> > > > Thanks, > >> > > > -ej > >> > > > > >> > > > On Wed, Aug 12, 2026 at 2:05 PM EJ Wang < > >> > [email protected]> > >> > > > wrote: > >> > > > > >> > > >> Thanks Dmitri, these comments were very useful. > >> > > >> > >> > > >> I went through the three areas you called out and updated the > >> proposal > >> > > >> accordingly. > >> > > >> > >> > > >> On the permission/policy direction, *I agree the Tag model should > >> > leave > >> > > >> room for permissions or policies to consume tags later*, > including > >> the > >> > > >> direction JB proposed. I am keeping that outside the v1 Tag > >> contract, > >> > > >> though. In v1, tags classify resources; they do not themselves > >> grant > >> > or > >> > > >> deny access. Polaris Policy looks like the closest existing > >> foundation > >> > > if > >> > > >> we later want a portable tag-aware policy model, but I think that > >> > > deserves > >> > > >> a separate proposal rather than baking policy semantics into the > >> Tag > >> > > >> storage model now. > >> > > >> > >> > > >> I also made the authorizer path more explicit. *A future OPA, > >> Ranger, > >> > or > >> > > >> other authorizer could receive the target's complete effective > >> tags as > >> > > >> resource attributes*. The authorization path would resolve those > >> tags > >> > > >> internally, applying target-types, inheritance, closest-wins, > >> > > grandfathered > >> > > >> values, and the same coherent-read guarantees as the Tag API. At > >> > > minimum, > >> > > >> the portable input can include the tag definition ID, current > name, > >> > and > >> > > >> selected value; provenance can be additional context. If Polaris > >> > cannot > >> > > >> resolve the complete effective state, authorization should fail > >> closed > >> > > >> rather than treat the resource as untagged. > >> > > >> > >> > > >> That also makes the persistence expectation on the read path > >> clearer: > >> > an > >> > > >> implementation needs to resolve the target and relevant > ancestors, > >> > > obtain > >> > > >> the applicable tag definitions and assignments, and produce one > >> > coherent > >> > > >> effective result. *Those observable semantics are the backend > >> > contract; > >> > > >> the physical lookup/indexing strategy is not.* > >> > > >> > >> > > >> On the Java interface suggestion, I added Java-shaped records for > >> the > >> > > >> durable logical model so the definition, target identity, and > >> > assignment > >> > > >> shapes are easier to review from JDBC and NoSQL perspectives. I > >> > stopped > >> > > >> short of proposing operation interfaces in pseudo-code, though. > My > >> > > current > >> > > >> thinking is that we should first agree on the durable facts and > >> > required > >> > > >> behavior, then design the actual persistence SPI around the needs > >> of > >> > the > >> > > >> implementations. I did not want an illustrative interface in this > >> > > design to > >> > > >> accidentally become the persistence contract. > >> > > >> > >> > > >> So Part 2 now separates the two intentionally: > >> > > >> > >> > > >> *logical data + behavior/conformance requirements are specified; > >> > > >> transaction, CAS, atomic batch, provider-native operations, and > the > >> > > >> eventual Java SPI remain implementation/design choices.* > >> > > >> > >> > > >> Thanks again for the review, and definitely keep the comments > >> coming > >> > :) > >> > > >> > >> > > >> I've updated the doc, please check it out the latest and the > >> greatest: > >> > > >> > >> > > >> > >> > > > >> > > >> > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0 > >> > > >> > >> > > >> -ej > >> > > >> > >> > > >> On Fri, Aug 7, 2026 at 3:39 PM Dmitri Bourlatchkov < > >> [email protected]> > >> > > >> wrote: > >> > > >> > >> > > >>> Hi EJ, JB, > >> > > >>> > >> > > >>> I left some comments on EJ's doc. I actually have a lot of > >> comments > >> > on > >> > > >>> the > >> > > >>> REST API design, I only posted some of them to start a > discussion > >> > > >>> without overloading the doc. > >> > > >>> > >> > > >>> Overall, I believe EJ's proposal should also allow permission > >> > > assignments > >> > > >>> on tags that JB proposed (eventually). We just need to clearly > >> define > >> > > the > >> > > >>> persistence expectations for looking up related tags on the read > >> > path. > >> > > >>> > >> > > >>> We should probably specify whether and how tags are exposed to > >> > > >>> authorizers > >> > > >>> (OPA, Ranger). I imagine people will want to use them in > external > >> > > policy > >> > > >>> engines the moment the feature is available. > >> > > >>> > >> > > >>> On the persistence side, I believe it would be nice to define > >> actual > >> > > java > >> > > >>> interfaces (perhaps in pseudo code) to allow easier review from > >> the > >> > > NoSQL > >> > > >>> persistence perspective (also commented in the doc). > >> > > >>> > >> > > >>> Cheers, > >> > > >>> Dmitri. > >> > > >>> > >> > > >>> On Thu, Jul 30, 2026 at 12:39 AM Jean-Baptiste Onofré < > >> > [email protected] > >> > > > > >> > > >>> wrote: > >> > > >>> > >> > > >>> > Hi EJ > >> > > >>> > > >> > > >>> > Thanks for starting this discussion. > >> > > >>> > > >> > > >>> > For the record, here's my initial proposal about tagging: > >> > > >>> > > >> https://lists.apache.org/thread/nmqmmjfmocfllb71fcmyp9syc9gyn820 > >> > > >>> > > >> > > >>> > At that time, only Dmitri replied :) > >> > > >>> > So, I would be happy to work with you on this, as I still have > >> the > >> > > PoC > >> > > >>> > I created for my initial proposal. > >> > > >>> > > >> > > >>> > I will try to join the scheduled meeting (no guarantee). > >> > > >>> > > >> > > >>> > Regards > >> > > >>> > JB > >> > > >>> > > >> > > >>> > On Fri, Jul 17, 2026 at 6:53 AM EJ Wang < > >> > > >>> [email protected]> > >> > > >>> > wrote: > >> > > >>> > > > >> > > >>> > > Hi folks, > >> > > >>> > > > >> > > >>> > > I have prepared a Google Doc > >> > > >>> > > < > >> > > >>> > > >> > > >>> > >> > > > >> > > >> > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing > >> > > >>> > > > >> > > >>> > > for the Polaris tag spec proposal. > >> > > >>> > > > >> > > >>> > > The goal is simple: add a native tag model to Polaris so > users > >> > can > >> > > >>> > classify > >> > > >>> > > catalog objects, read those classifications back, and find > >> > objects > >> > > by > >> > > >>> > tag. > >> > > >>> > > > >> > > >>> > > The proposal covers: > >> > > >>> > > * tag definitions as catalog-scoped Polaris entities > >> > > >>> > > * tag assignments on catalogs, namespaces, table-like > objects, > >> > and > >> > > >>> > columns > >> > > >>> > > * allowed values on tag definitions > >> > > >>> > > * direct and inherited tag reads > >> > > >>> > > * direct by-tag lookup > >> > > >>> > > * the durable model behind the API > >> > > >>> > > * how this compares with the existing Polaris Policy API > (tag > >> > > design > >> > > >>> > > referenced policy heavily, given their pattern similarity) > >> > > >>> > > > >> > > >>> > > Please take a look and leave comments in the doc. Let me > know > >> > WDYT! > >> > > >>> > > > >> > > >>> > > I would also like to discuss this in the July 23 community > >> sync. > >> > A > >> > > >>> > separate > >> > > >>> > > dedicated review meeting will be scheduled separately, > likely > >> > > within > >> > > >>> the > >> > > >>> > > next two weeks. > >> > > >>> > > > >> > > >>> > > Thanks, > >> > > >>> > > -ej > >> > > >>> > > >> > > >>> > >> > > >> > >> > > > >> > > >> > > >
