Hi EJ, On pagination:
> [...] I'd still like to preserve an explicit pagination=false option for > users accustomed to IRC's full-result mode [...] > That option remains subject to finite server limits on response size and > work. It returns the complete result or an error, never a silently > truncated result. Requests exceeding those limits must use pagination. This sounds reasonable to me. Cheers, Dmitri. On Mon, Sep 21, 2026 at 7:34 PM EJ Wang <[email protected]> wrote: > Hi Dmitri, all, > > I've updated the design doc, and the corresponding PR changes are in > progress: > > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit > > On the API root, I see the benefits of independent routing and versioning, > and agree that build wiring shouldn’t decide the paths. Would you be > comfortable letting PR1 proceed under the current catalog root while we > continue discussing the broader API layout? There are still implementation > PRs ahead, so we can revisit this before release. If a later decision > requires migration, we could introduce the new root while retaining the > current paths for compatibility, accepting the maintenance cost. > > On pagination, the doc now defaults to paginated responses, with a > server-selected page size when omitted and client-provided sizes treated as > hints. I understand that your oversized-response workaround was > specifically for IRC compatibility. Tags has no existing-client > compatibility requirement, but I'd still like to preserve an explicit > pagination=false option for users accustomed to IRC's full-result mode as > they adapt their workflows to paginated defaults. > > That option remains subject to finite server limits on response size and > work. It returns the complete result or an error, never a silently > truncated result. Requests exceeding those limits must use pagination. > Would this bounded opt-out be acceptable for Tags? > > Two corrections to my earlier replies to Robert and Prithvi: reverse lookup > now requires per-target property-read checks, and detach-all requires the > same atomic visible result from every implementation supporting the API. > Physical cleanup may follow, but the spec no longer defines a 501 fallback > for that guarantee. > > -ej > > On Fri, Sep 18, 2026 at 1:31 PM Dmitri Bourlatchkov <[email protected]> > wrote: > > > Hi EJ, > > > > > 1. *Base URL*: *existing catalog root* vs. *independent Tag root* > > > > Forwarding my preference for a separate URI base path (option B). I > believe > > I commented about that in GH too. > > > > While each endpoint in the proposed Tags API is distinct from endpoints > > currently under /api/catalog/polaris/v1, I do not think interleaving > > endpoints from conceptually different APIs under the same URI base is a > > good idea. > > > > If API under /api/catalog/polaris/v1 were formulated in a way to allow > > exencibility, then adding tags there would be natural. However, I do not > > think existing endpoints under that base URI were designed for > > extencibility. The form very narrow purpose API. Therefore, I believe it > is > > preferable to put tags under a separate root. > > > > A smaller point is API evolution. Tags are defined in a separate Open API > > YAML file. That scopes the API down and naturally isolates it in terms of > > API evolution. > > > > Making changes to the tags API will require careful consideration of the > > other APIs under the same base URI to ensure no overlaps. > > > > That said, I do not feel too strongly about the base URI. > > > > > 2. *Pagination*: *opt-in* vs. *always-on* > > > > Forcing servers to provide unpaginated responses for potentially large > > datasets is a DoS / overload risk, IMHO. More in-depth discussion on this > > is in [1]. > > > > I would not want the Polaris API specs to repeat that guideline from the > > IRC spec as I think it is fundamentally flawed. > > > > I previously suggested [2] failing large responses only as a means for > > maintaining IRC spec compatibility in the IRC API. > > > > So I propose pagination to be "always on" in the API spec. So, all > clients > > should be prepared to handle paginated responses. > > > > However, pagination on the server side can be implemented in phases, if > it > > helps with code-level PRs. > > > > [1] https://lists.apache.org/thread/k81ptyktdbdf8gynncgk3o04mqt85zyk > > > > [2] https://lists.apache.org/thread/ntn71oh7g0kkf1wdhskh8t986gdwh1p3 > > > > Cheers, > > Dmitri > > > > On Tue, Sep 15, 2026 at 6:46 PM EJ Wang <[email protected]> > > wrote: > > > > > Hi folks, > > > > > > Following Dmitri's review of PR1, I'd like community input on two > choices > > > before updating the spec. > > > > > > 1. *Base URL*: *existing catalog root* vs. *independent Tag root* > > > > > > *A. Keep /api/catalog/polaris/v1/{prefix}/.* Follow the existing policy > > and > > > generic-table APIs, and discuss the broader extension API layout > > > separately. > > > > > > *B. Move to /api/tags/v1/{prefix}/.* Give Tags independent versioning > and > > > routing from its first release. Other APIs would not need to move in > this > > > PR. > > > > > > *My preference is A*, revising my earlier agreement to B in the PR. I > see > > > the benefits of an independent root, but currently favor consistency > with > > > the existing layout over introducing another pattern for Tags alone. > > > > > > 2. *Pagination*: *opt-in* vs. *always-on* > > > > > > *A. Opt-in*. Omitting pageToken requests all results in one response. > An > > > empty token starts pagination. This follows Iceberg's convention and > lets > > > simple clients avoid implementing pagination. > > > *B. Always-on*. Omitting pagination parameters returns a default-sized > > > first page. Clients must follow continuation tokens to retrieve all > > > results. This is Dmitri's proposal and lets servers bound response > sizes. > > > > > > *I'd like to explore A with an explicit safeguard*: oversized > unpaginated > > > requests could fail with a documented error, never silently truncate > > > results. That adds a limit to the full-result behavior, so it would > need > > to > > > be part of the contract. > > > > > > The main concern is reverse lookup, where a tag may reference many > > objects. > > > Internal batching would not bound the total response size or request > > > duration. The argument for A is client simplicity and convention > > > consistency, not compatibility with existing Tag clients. > > > > > > Which option do you favor for each? For pagination, would the > > > explicit-error safeguard make A acceptable, or is B preferable from the > > > start? > > > > > > Thanks, > > > -ej > > > > > > On Fri, Sep 11, 2026 at 7:55 AM Dmitri Bourlatchkov <[email protected]> > > > wrote: > > > > > > > Hi All, > > > > > > > > I posted a comment about tags API URI paths that may be of interest > to > > a > > > > few people. Re-posting here for awareness: > > > > > > > > https://github.com/apache/polaris/pull/5366#discussion_r3980198150 > > > > > > > > Cheers, > > > > Dmitri. > > > > > > > > On Thu, Sep 10, 2026 at 1:21 AM EJ Wang < > > [email protected]> > > > > wrote: > > > > > > > > > Hi folks, > > > > > > > > > > A quick update on the Tag work. PR1-3 and the design doc have been > > > > updated > > > > > following the discussion. > > > > > > > > > > *Current status*: > > > > > PR1: API contract: ready for review. > > > > > PR2: Definition CRUD: open as Draft, ready for review if PR1 LGTY. > > > > > *NEW! *PR3: Assignment writes and storage: open as Draft, ready for > > > > review > > > > > if PR2 LGTY. > > > > > PR4: Reads, inheritance, and reverse lookup: planned. > > > > > > > > > > *More on PR3*: > > > > > PR3 adds assign/unassign for catalogs, namespaces, tables, and > > > top-level > > > > > Iceberg columns, plus atomic detach-all. It includes JDBC and > > in-memory > > > > > support. Reads remain in PR4. > > > > > The PR is stacked on PR2. Its description links the assignment-only > > > > commit > > > > > for focused review. > > > > > > > > > > *Asks*: > > > > > For PR3, I’d especially appreciate feedback on the persistence > > > > interfaces, > > > > > their impact on external implementations, and the detach-all > > guarantee. > > > > > The updated design doc also includes the reverse-lookup flow and a > > > > > discussion of the deferred shared encoding work. > > > > > > > > > > Thanks, > > > > > -ej > > > > > > > > > > On Wed, Sep 9, 2026 at 3:09 PM EJ Wang < > > [email protected] > > > > > > > > > wrote: > > > > > > > > > > > Thanks Robert, Prithvi, and Dmitri. I’ve updated the spec doc > > > > > > < > > > > > > > > > > > > > > > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0#heading=h.jx650zq2bd2g > > > > > > > > > > with > > > > > > the contract clarifications (see below) and the follow-up > > discussion > > > > with > > > > > > Dmitri. The corresponding PR updates are underway. > > > > > > > > > > > > *For encoding*, my proposal is to retain Iceberg’s namespace > query > > > > > > convention in v1, with explicit supported-name, encoding, and > > > decoding > > > > > > rules in Part 1 section 5.2. This retains the known separator and > > > > ingress > > > > > > limitations. Since the issue also affects existing > Iceberg/Polaris > > > > APIs, > > > > > > I’d address a replacement codec in a separate issue/PR. Part 3 > > > section > > > > > 7.12 > > > > > > records alternatives and spike results to start that discussion. > > > > > > > > > > > > *Version tokens* are opaque strings in both responses and update > > > > > > requests. Clients return the token unchanged, and stale updates > > > return > > > > > 409. > > > > > > The check must cover every supported definition-write path, while > > > > > backends > > > > > > choose their revision mechanism. > > > > > > > > > > > > *For detach-all*, the definition and assignments must disappear > > > > together > > > > > > through Tag reads, or nothing changes. Physical cleanup may > follow. > > > An > > > > > > implementation unable to provide that guarantee returns 501 after > > > > > > authorization and before changing visible state. > > > > > > > > > > > > *Column assignments* use table identity and column identifier, > > using > > > > > > Iceberg field id for Iceberg tables. Renames preserve assignments > > > when > > > > > > identity is preserved, while same-name replacements do not > inherit > > > > them. > > > > > V1 > > > > > > remains limited to top-level Iceberg columns. > > > > > > > > > > > > Does this scope split work for you, particularly keeping the > shared > > > > codec > > > > > > redesign separate from Tag v1? I’d like to settle these contract > > > points > > > > > > here before PR1 merges. > > > > > > > > > > > > -ej > > > > > > > > > > > > On Wed, Sep 9, 2026 at 7:36 AM Dmitri Bourlatchkov < > > [email protected] > > > > > > > > > > wrote: > > > > > > > > > > > >> Hi All, > > > > > >> > > > > > >> (replying partially) > > > > > >> > > > > > >> I very much support Robert's proposal for using a well-defined > and > > > > > >> unambiguous format for namespaces. > > > > > >> > > > > > >> Many Polaris APIs fall into following the IRC approach to > > namespace > > > > > >> representation in query parameters. Yet, that approach has > > multiple > > > > > issues > > > > > >> , which can be seen in Iceberg dev ML / GH issues. > > > > > >> > > > > > >> I think Polaris should use a more robust namespace > representation > > in > > > > its > > > > > >> native APIs. > > > > > >> > > > > > >> Cheers, > > > > > >> Dmitri. > > > > > >> > > > > > >> On Tue, Sep 8, 2026 at 6:41 AM Robert Stupp <[email protected]> > > wrote: > > > > > >> > > > > > >> > Hi, > > > > > >> > > > > > > >> > I have been thinking about what a full implementation would > > > require. > > > > > >> > > > > > > >> > I like that the proposal separates definitions, assignments, > and > > > > > >> effective > > > > > >> > reads. I have a few API-contract questions that seem worth > > > resolving > > > > > >> while > > > > > >> > the contract is still separate from the implementation. > > > > > >> > > > > > > >> > > > > > > >> > First, I think the target query parameters need a defined > > encoding > > > > for > > > > > >> > identifier elements. > > > > > >> > > > > > > >> > This is not only about the unit separator. Namespace elements > > and > > > > > object > > > > > >> > names may themselves contain characters such as &, ?, =, +, or > > %. > > > > > >> > Those must be preserved rather than interpreted as query > syntax. > > > > > >> > > > > > > >> > Ordinary URI query-value encoding handles those characters, > but > > it > > > > > does > > > > > >> not > > > > > >> > solve the separate problem of representing the boundaries > > between > > > > > >> multipart > > > > > >> > namespace elements. > > > > > >> > > > > > > >> > I suggest defining a small, reversible namespace-element > codec, > > > then > > > > > >> > applying ordinary URI encoding to its complete output. > Nessie's > > > > > escaped > > > > > >> > path > > > > > >> > representation is a useful precedent: it has an unambiguous > > > element > > > > > >> > separator > > > > > >> > and escape syntax, while avoiding control characters in the > > > > transport > > > > > >> > representation. > > > > > >> > > > > > > >> > The contract should specify the codec, its decoding failures, > > and > > > > > >> > conformance examples. Client libraries should expose it rather > > > than > > > > > >> > requiring > > > > > >> > every client to reproduce it. The structured target used by > the > > > > write > > > > > >> APIs > > > > > >> > would still be the clearest canonical representation; this > codec > > > > would > > > > > >> make > > > > > >> > the GET form safe and interoperable. > > > > > >> > > > > > > >> > > > > > > >> > Second, I think the revision token should be opaque at the API > > > > > boundary. > > > > > >> > > > > > > >> > The backend should be free to use a native row revision, > commit > > > ID, > > > > > >> ETag, > > > > > >> > or > > > > > >> > another conditional-write token. However, the contract should > > > define > > > > > the > > > > > >> > observable precondition: the server returns a token, and an > > update > > > > > >> succeeds > > > > > >> > only if the client supplies the token for the current tag > > > > definition. > > > > > >> > Otherwise the server returns a conflict. > > > > > >> > > > > > > >> > That requires token matching semantics, but not an integer > type, > > > an > > > > > >> initial > > > > > >> > value, ordering, increment-by-one behavior, or history > > semantics. > > > A > > > > > >> > catalog-wide commit token would also be valid, although it > could > > > > > create > > > > > >> > avoidable conflicts for unrelated changes. > > > > > >> > > > > > > >> > > > > > > >> > Third, the direct reverse lookup is useful, but I would treat > it > > > as > > > > a > > > > > >> > first-class, paginated relationship rather than a tag record > > > > > containing > > > > > >> a > > > > > >> > collection of targets. > > > > > >> > > > > > > >> > A common tag can legitimately be attached to a very large > number > > > of > > > > > >> objects > > > > > >> > or columns. A backend will normally need one forward access > path > > > for > > > > > >> direct > > > > > >> > assignments by target, and one reverse access path by > tag/value, > > > > with > > > > > >> > backend-specific partitioning or sharding. Effective > assignments > > > > > should > > > > > >> > remain > > > > > >> > computed from the target and its ancestors; materializing > > > inherited > > > > > >> > assignments onto descendants would have very different scaling > > > > > behavior. > > > > > >> > > > > > > >> > This also affects detach-all. Deleting an unbounded number of > > > > > assignment > > > > > >> > records atomically is not a portable primitive for all > backends. > > > The > > > > > >> > contract > > > > > >> > should distinguish observable deletion semantics from physical > > > > > cleanup, > > > > > >> or > > > > > >> > state the backend capability required for a synchronous > > detach-all > > > > > >> > operation. > > > > > >> > > > > > > >> > > > > > > >> > Finally, I agree with the V1 boundary of top-level Iceberg > > > columns, > > > > > but > > > > > >> I > > > > > >> > would > > > > > >> > keep the core tag model independent of Iceberg. For Iceberg, > the > > > > > durable > > > > > >> > column reference should be the field ID, with a column name > used > > > > only > > > > > >> for > > > > > >> > request-time resolution and display. Other table > implementations > > > > could > > > > > >> opt > > > > > >> > in > > > > > >> > later once they provide an equally stable field identity. This > > > > avoids > > > > > >> > treating > > > > > >> > a name-based column mapping as a general abstraction. > > > > > >> > > > > > > >> > > > > > > >> > None of this requires tags to become authorization inputs in > V1. > > > It > > > > is > > > > > >> > mainly > > > > > >> > about leaving the assignment and read contract implementable > by > > > more > > > > > >> than > > > > > >> > one > > > > > >> > persistence model when those slices arrive. > > > > > >> > > > > > > >> > Thanks, > > > > > >> > Robert > > > > > >> > > > > > > >> > > > > > > >> > On Fri, Aug 28, 2026 at 2:19 AM EJ Wang < > > > > > [email protected] > > > > > >> > > > > > > >> > wrote: > > > > > >> > > > > > > >> > > Hi folks, > > > > > >> > > > > > > > >> > > A quick update on the Tag work. *Current status*: > > > > > >> > > - PR1: API contract ( > > > https://github.com/apache/polaris/pull/5366 > > > > ): > > > > > >> ready > > > > > >> > > for review > > > > > >> > > - *NEW! *PR2: Definition CRUD ( > > > > > >> > https://github.com/apache/polaris/pull/5391 > > > > > >> > > ): > > > > > >> > > open as Draft, ready for review if PR1 LGTY > > > > > >> > > - PR3: Assignment writes and storage: planned > > > > > >> > > - PR4: Reads, inheritance, and reverse lookup: planned > > > > > >> > > > > > > > >> > > *More on PR2:* > > > > > >> > > - PR2 makes Tag definitions usable through create, list, > load, > > > > > update, > > > > > >> > > rename, and delete. It intentionally stops before > assignments, > > > so > > > > > >> their > > > > > >> > > persistence model remains open for the next slice. > > > > > >> > > - The PR is stacked on #5366 and will be rebased once that > PR > > > > > merges. > > > > > >> > > > > > > > >> > > *Asks:* > > > > > >> > > - For #5366, please call out any remaining API contract > > > concerns. > > > > > For > > > > > >> > > #5391, I would especially appreciate feedback on the slice > > > > boundary > > > > > >> and > > > > > >> > the > > > > > >> > > decision to reuse the existing entity persistence model. > > > > > >> > > - The design doc remains here: > > > > > >> > > > > > > > >> > > > > > > > >> > > > > > > >> > > > > > > > > > > > > > > > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing > > > > > >> > > > > > > > >> > > I’ll keep using this thread for new delivery slices, > material > > > > status > > > > > >> > > changes, and specific community asks. > > > > > >> > > > > > > > >> > > Thanks, > > > > > >> > > -ej > > > > > >> > > > > > > > >> > > On Mon, Aug 24, 2026 at 5:04 PM EJ Wang < > > > > > >> [email protected]> > > > > > >> > > wrote: > > > > > >> > > > > > > > >> > > > Hi folks, > > > > > >> > > > > > > > > >> > > > Following up on this thread, I have opened a PR to land > the > > > > public > > > > > >> API > > > > > >> > > > contract for Tags: > > > https://github.com/apache/polaris/pull/5366 > > > > > >> > > > > > > > > >> > > > The PR defines Tag management, assignment and > unassignment, > > > > direct > > > > > >> and > > > > > >> > > > inherited reads, and reverse lookup. V1 covers catalogs, > > > > > namespaces, > > > > > >> > > > Iceberg and generic tables as whole objects, and top-level > > > > Iceberg > > > > > >> > table > > > > > >> > > > columns. Views, generic-table columns, nested fields, > > > > multi-value > > > > > >> > > > assignments, and tag-based authorization are deferred. > > > > > >> > > > > > > > > >> > > > I plan to deliver the capability through four PRs that > merge > > > in > > > > > >> order: > > > > > >> > > the > > > > > >> > > > API contract in this PR, Tag definition CRUD, assignment > > > writes > > > > > and > > > > > >> > > > storage, then reads and reverse lookup. A separate > follow-up > > > > will > > > > > >> add > > > > > >> > > > grants on Tag resources to the management APIs. That grant > > > > surface > > > > > >> is > > > > > >> > > > distinct from using Tags to control access to tagged > > objects, > > > > > which > > > > > >> > > remains > > > > > >> > > > outside v1. > > > > > >> > > > > > > > > >> > > > The updated design doc is here: > > > > > >> > > > > > > > > >> > > > > > > > >> > > > > > > >> > > > > > > > > > > > > > > > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing > > > > > >> > > > > > > > > >> > > > The PR is currently Draft while we finish aligning on the > > > public > > > > > >> > > contract. > > > > > >> > > > It is intended to merge as the first delivery slice, not > > > remain > > > > > as a > > > > > >> > > > design-only artifact. Please call out any remaining scope > or > > > > > >> contract > > > > > >> > > > concerns. If the list is aligned, I will mark it ready for > > > > review. > > > > > >> > > > > > > > > >> > > > Thanks, > > > > > >> > > > -ej > > > > > >> > > > > > > > > >> > > > On Wed, Aug 12, 2026 at 2:05 PM EJ Wang < > > > > > >> > [email protected]> > > > > > >> > > > wrote: > > > > > >> > > > > > > > > >> > > >> Thanks Dmitri, these comments were very useful. > > > > > >> > > >> > > > > > >> > > >> I went through the three areas you called out and updated > > the > > > > > >> proposal > > > > > >> > > >> accordingly. > > > > > >> > > >> > > > > > >> > > >> On the permission/policy direction, *I agree the Tag > model > > > > should > > > > > >> > leave > > > > > >> > > >> room for permissions or policies to consume tags later*, > > > > > including > > > > > >> the > > > > > >> > > >> direction JB proposed. I am keeping that outside the v1 > Tag > > > > > >> contract, > > > > > >> > > >> though. In v1, tags classify resources; they do not > > > themselves > > > > > >> grant > > > > > >> > or > > > > > >> > > >> deny access. Polaris Policy looks like the closest > existing > > > > > >> foundation > > > > > >> > > if > > > > > >> > > >> we later want a portable tag-aware policy model, but I > > think > > > > that > > > > > >> > > deserves > > > > > >> > > >> a separate proposal rather than baking policy semantics > > into > > > > the > > > > > >> Tag > > > > > >> > > >> storage model now. > > > > > >> > > >> > > > > > >> > > >> I also made the authorizer path more explicit. *A future > > OPA, > > > > > >> Ranger, > > > > > >> > or > > > > > >> > > >> other authorizer could receive the target's complete > > > effective > > > > > >> tags as > > > > > >> > > >> resource attributes*. The authorization path would > resolve > > > > those > > > > > >> tags > > > > > >> > > >> internally, applying target-types, inheritance, > > closest-wins, > > > > > >> > > grandfathered > > > > > >> > > >> values, and the same coherent-read guarantees as the Tag > > API. > > > > At > > > > > >> > > minimum, > > > > > >> > > >> the portable input can include the tag definition ID, > > current > > > > > name, > > > > > >> > and > > > > > >> > > >> selected value; provenance can be additional context. If > > > > Polaris > > > > > >> > cannot > > > > > >> > > >> resolve the complete effective state, authorization > should > > > fail > > > > > >> closed > > > > > >> > > >> rather than treat the resource as untagged. > > > > > >> > > >> > > > > > >> > > >> That also makes the persistence expectation on the read > > path > > > > > >> clearer: > > > > > >> > an > > > > > >> > > >> implementation needs to resolve the target and relevant > > > > > ancestors, > > > > > >> > > obtain > > > > > >> > > >> the applicable tag definitions and assignments, and > produce > > > one > > > > > >> > coherent > > > > > >> > > >> effective result. *Those observable semantics are the > > backend > > > > > >> > contract; > > > > > >> > > >> the physical lookup/indexing strategy is not.* > > > > > >> > > >> > > > > > >> > > >> On the Java interface suggestion, I added Java-shaped > > records > > > > for > > > > > >> the > > > > > >> > > >> durable logical model so the definition, target identity, > > and > > > > > >> > assignment > > > > > >> > > >> shapes are easier to review from JDBC and NoSQL > > > perspectives. I > > > > > >> > stopped > > > > > >> > > >> short of proposing operation interfaces in pseudo-code, > > > though. > > > > > My > > > > > >> > > current > > > > > >> > > >> thinking is that we should first agree on the durable > facts > > > and > > > > > >> > required > > > > > >> > > >> behavior, then design the actual persistence SPI around > the > > > > needs > > > > > >> of > > > > > >> > the > > > > > >> > > >> implementations. I did not want an illustrative interface > > in > > > > this > > > > > >> > > design to > > > > > >> > > >> accidentally become the persistence contract. > > > > > >> > > >> > > > > > >> > > >> So Part 2 now separates the two intentionally: > > > > > >> > > >> > > > > > >> > > >> *logical data + behavior/conformance requirements are > > > > specified; > > > > > >> > > >> transaction, CAS, atomic batch, provider-native > operations, > > > and > > > > > the > > > > > >> > > >> eventual Java SPI remain implementation/design choices.* > > > > > >> > > >> > > > > > >> > > >> Thanks again for the review, and definitely keep the > > comments > > > > > >> coming > > > > > >> > :) > > > > > >> > > >> > > > > > >> > > >> I've updated the doc, please check it out the latest and > > the > > > > > >> greatest: > > > > > >> > > >> > > > > > >> > > >> > > > > > >> > > > > > > > >> > > > > > > >> > > > > > > > > > > > > > > > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?pli=1&tab=t.0 > > > > > >> > > >> > > > > > >> > > >> -ej > > > > > >> > > >> > > > > > >> > > >> On Fri, Aug 7, 2026 at 3:39 PM Dmitri Bourlatchkov < > > > > > >> [email protected]> > > > > > >> > > >> wrote: > > > > > >> > > >> > > > > > >> > > >>> Hi EJ, JB, > > > > > >> > > >>> > > > > > >> > > >>> I left some comments on EJ's doc. I actually have a lot > > of > > > > > >> comments > > > > > >> > on > > > > > >> > > >>> the > > > > > >> > > >>> REST API design, I only posted some of them to start a > > > > > discussion > > > > > >> > > >>> without overloading the doc. > > > > > >> > > >>> > > > > > >> > > >>> Overall, I believe EJ's proposal should also allow > > > permission > > > > > >> > > assignments > > > > > >> > > >>> on tags that JB proposed (eventually). We just need to > > > clearly > > > > > >> define > > > > > >> > > the > > > > > >> > > >>> persistence expectations for looking up related tags on > > the > > > > read > > > > > >> > path. > > > > > >> > > >>> > > > > > >> > > >>> We should probably specify whether and how tags are > > exposed > > > to > > > > > >> > > >>> authorizers > > > > > >> > > >>> (OPA, Ranger). I imagine people will want to use them in > > > > > external > > > > > >> > > policy > > > > > >> > > >>> engines the moment the feature is available. > > > > > >> > > >>> > > > > > >> > > >>> On the persistence side, I believe it would be nice to > > > define > > > > > >> actual > > > > > >> > > java > > > > > >> > > >>> interfaces (perhaps in pseudo code) to allow easier > review > > > > from > > > > > >> the > > > > > >> > > NoSQL > > > > > >> > > >>> persistence perspective (also commented in the doc). > > > > > >> > > >>> > > > > > >> > > >>> Cheers, > > > > > >> > > >>> Dmitri. > > > > > >> > > >>> > > > > > >> > > >>> On Thu, Jul 30, 2026 at 12:39 AM Jean-Baptiste Onofré < > > > > > >> > [email protected] > > > > > >> > > > > > > > > >> > > >>> wrote: > > > > > >> > > >>> > > > > > >> > > >>> > Hi EJ > > > > > >> > > >>> > > > > > > >> > > >>> > Thanks for starting this discussion. > > > > > >> > > >>> > > > > > > >> > > >>> > For the record, here's my initial proposal about > > tagging: > > > > > >> > > >>> > > > > > > >> > https://lists.apache.org/thread/nmqmmjfmocfllb71fcmyp9syc9gyn820 > > > > > >> > > >>> > > > > > > >> > > >>> > At that time, only Dmitri replied :) > > > > > >> > > >>> > So, I would be happy to work with you on this, as I > > still > > > > have > > > > > >> the > > > > > >> > > PoC > > > > > >> > > >>> > I created for my initial proposal. > > > > > >> > > >>> > > > > > > >> > > >>> > I will try to join the scheduled meeting (no > guarantee). > > > > > >> > > >>> > > > > > > >> > > >>> > Regards > > > > > >> > > >>> > JB > > > > > >> > > >>> > > > > > > >> > > >>> > On Fri, Jul 17, 2026 at 6:53 AM EJ Wang < > > > > > >> > > >>> [email protected]> > > > > > >> > > >>> > wrote: > > > > > >> > > >>> > > > > > > > >> > > >>> > > Hi folks, > > > > > >> > > >>> > > > > > > > >> > > >>> > > I have prepared a Google Doc > > > > > >> > > >>> > > < > > > > > >> > > >>> > > > > > > >> > > >>> > > > > > >> > > > > > > > >> > > > > > > >> > > > > > > > > > > > > > > > https://docs.google.com/document/d/1rIJGzcsmGhfrBiRXPac51hr-jeJuuKQQBYjgBdOb9-k/edit?usp=sharing > > > > > >> > > >>> > > > > > > > >> > > >>> > > for the Polaris tag spec proposal. > > > > > >> > > >>> > > > > > > > >> > > >>> > > The goal is simple: add a native tag model to > Polaris > > so > > > > > users > > > > > >> > can > > > > > >> > > >>> > classify > > > > > >> > > >>> > > catalog objects, read those classifications back, > and > > > find > > > > > >> > objects > > > > > >> > > by > > > > > >> > > >>> > tag. > > > > > >> > > >>> > > > > > > > >> > > >>> > > The proposal covers: > > > > > >> > > >>> > > * tag definitions as catalog-scoped Polaris entities > > > > > >> > > >>> > > * tag assignments on catalogs, namespaces, > table-like > > > > > objects, > > > > > >> > and > > > > > >> > > >>> > columns > > > > > >> > > >>> > > * allowed values on tag definitions > > > > > >> > > >>> > > * direct and inherited tag reads > > > > > >> > > >>> > > * direct by-tag lookup > > > > > >> > > >>> > > * the durable model behind the API > > > > > >> > > >>> > > * how this compares with the existing Polaris Policy > > API > > > > > (tag > > > > > >> > > design > > > > > >> > > >>> > > referenced policy heavily, given their pattern > > > similarity) > > > > > >> > > >>> > > > > > > > >> > > >>> > > Please take a look and leave comments in the doc. > Let > > me > > > > > know > > > > > >> > WDYT! > > > > > >> > > >>> > > > > > > > >> > > >>> > > I would also like to discuss this in the July 23 > > > community > > > > > >> sync. > > > > > >> > A > > > > > >> > > >>> > separate > > > > > >> > > >>> > > dedicated review meeting will be scheduled > separately, > > > > > likely > > > > > >> > > within > > > > > >> > > >>> the > > > > > >> > > >>> > > next two weeks. > > > > > >> > > >>> > > > > > > > >> > > >>> > > Thanks, > > > > > >> > > >>> > > -ej > > > > > >> > > >>> > > > > > > >> > > >>> > > > > > >> > > >> > > > > > >> > > > > > > > >> > > > > > > >> > > > > > > > > > > > > > > > > > > > > >
