Hi Andrei, Could you please elaborate on what issues you are trying to solve in iceberg-go with a labels metadata table?
Andrei Tserakhau via dev <[email protected]>于2026年10月3日 周六07:17写道: > Hi all, > > Reviving this with a narrower question after the same issue came up again > in iceberg-go [1]. > > The java labels metadata-table PR [0] was closed because metadata tables > in Java/Spark are expected to be deterministic projections of table > metadata. I think that still makes sense for the engine SQL / Spark surface. > > The non-Java clients are a bit different, though. In PyIceberg, > iceberg-rust, and iceberg-go, inspect is a library read API, not a SQL > metadata table, and there is no convinient alternative. > > My suggestion is to keep the Java/Spark decision as-is, but let each > client expose catalog-provided data through its inspect API if useful. For > labels, that means documenting it as catalog-provided, captured at load > time, and empty when the catalog returns none. > > If that split sounds reasonable, I'll proceed with the iceberg-go PR [1] > and mirror it in rust and python. > > Any objections to treating the client inspect API separately from the > engine metadata-table contract? > > Thanks, > Andrei > > [0] https://github.com/apache/iceberg/pull/18048 (Java core labels > metadata table, closed) > [1] https://github.com/apache/iceberg-go/pull/2101 (iceberg-go labels > inspect table) > > On Tue, Sep 29, 2026 at 12:19 AM Andrei Tserakhau < > [email protected]> wrote: > >> Fair - I think this brings the focus back to the main question: why do we >> need >> SQL at all. >> >> My answer is that there's a class of consumers whose only interface is >> SQL - >> SQL-native discovery / governance tooling, and most notably LLM agents >> that >> explore a warehouse through SQL. For them a Java capability >> (SupportsLabels) is >> not a thing; they can only reason over what they can query. >> >> What can cover that is `DESCRIBE` - #18049 surfaces both object and field >> labels >> there (labels.object.*, labels.field.<id>.*). So if there's no urge to >> join >> stuff, I see the point - we can keep it simple and not do the metadata >> table. >> >> We can revisit this topic later, once we have more data points and use >> cases, >> but for now I agree - it's not needed. >> >> Best, >> Andrei >> >> On Mon, Sep 28, 2026 at 11:30 PM Ryan Blue <[email protected]> wrote: >> >>> I don't agree with the composition argument. The example query you >>> provided doesn't make sense because there is no ON clause so you end up >>> with a cartesian join. Luckily, since you're looking for a specific label, >>> "owner", you end up with just one label and will aggregate the entire set >>> of files. So you end up with a result that looks reasonable, but you're >>> really just running two unrelated queries here: one to aggregate the total >>> size of live files in the table, and one to select the owner. >>> >>> I think that means that the only use case here is to expose this data to >>> users. But I think that this reasoning is that we need a metadata table >>> because we need SQL interaction and we need SQL interaction because . . . ? >>> It's a nice-to-have, sure, but I'm not convinced that anyone would miss it >>> if we didn't expose this directly to users. >>> >>> On Fri, Sep 25, 2026 at 2:40 PM Andrei Tserakhau via dev < >>> [email protected]> wrote: >>> >>>> Hi Ryan, >>>> >>>> Agree here: the engine-facing consumption (cost attribution, policy >>>> attachment) goes through SupportsLabels, no table needed. The table is for >>>> the other consumer - SQL/people or AI Agent :). Some cases that i see here: >>>> >>>> 1) Exploration. "which columns are classified as X here" is a query a >>>> person runs, not something an engine surfaces: >>>> >>>> SELECT field_name, key, value >>>> FROM prod.db.orders.labels >>>> WHERE scope = 'field' AND key = 'classification'; >>>> >>>> 2) Composition. .labels joins with .files / .partitions / .snapshots in >>>> one query - e.g. attribute bytes to an owner label, i.e. cost attribution >>>> as a report someone runs, not a log an engine emits: >>>> >>>> SELECT l.value AS owner, SUM(f.file_size_in_bytes) AS bytes >>>> FROM prod.db.orders.files f, prod.db.orders.labels l >>>> WHERE l.scope = 'object' AND l.key = 'owner' >>>> GROUP BY l.value; >>>> >>>> I think the key value to have a metadata table is joinability, you >>>> can’t have it with a programmatic label API. >>>> >>>> So there is a place for the SQL consumer, the engine path is a >>>> different usecase and the co-live together. >>>> >>>> Thanks, >>>> Andrei >>>> >>>> >>>> On Fri, Sep 25, 2026 at 10:51 PM Ryan Blue <[email protected]> wrote: >>>> >>>>> > Are we OK with catalog-provided metadata tables as a separate >>>>> category? >>>>> >>>>> I'm okay with providing metadata through a system table like this, as >>>>> long as we think that people will want to access this data that way. >>>>> Question 3, "If no, what should the SQL surface for labels be instead?" >>>>> makes me think that a SQL surface is _assumed_ to be needed. >>>>> >>>>> I don't think it is necessarily the case that we need to expose these >>>>> for SQL users. I thought that we wanted labels to expose additional >>>>> context >>>>> to engines for things like cost attribution logs or attaching an engine's >>>>> policy to a table. That doesn't require a table-like user surface. >>>>> >>>>> I'm fine adding a metadata table if there's a use for it, but if we >>>>> don't need one then it's simpler not to add and maintain it. And that >>>>> avoids needing to answer questions like this as well. >>>>> >>>>> >>>>> >>>>> On Fri, Sep 25, 2026 at 1:13 PM Andrei Tserakhau via dev < >>>>> [email protected]> wrote: >>>>> >>>>>> Hi all, >>>>>> >>>>>> Labels in the REST spec recently landed [1]. A catalog can now expose >>>>>> object-level and per-field labels on load-table responses. As one of the >>>>>> follow-ups, there is a proposal to add a .labels metadata table >>>>>> backed by catalog data. >>>>>> >>>>>> During discussion of the follow-ups ([3], [4]), Peter raised a good >>>>>> question on [3]: every metadata table today is derived from table >>>>>> metadata. >>>>>> .labels would be different because the data comes from the catalog, >>>>>> may vary by catalog, and may be absent if the catalog has nothing to >>>>>> return. >>>>>> >>>>>> I think this is less about Labels itself and more about a new *kind* >>>>>> of data we would expose: metadata owned by the catalog rather than by >>>>>> storage. Labels would be the first example, but later the same pattern >>>>>> could be used to expose other catalog information as metadata tables. >>>>>> >>>>>> So the broader question is: do we want metadata tables to also expose >>>>>> catalog-provided information? >>>>>> >>>>>> I think labels are a reasonable first case. They are structured, >>>>>> useful to query via SQL, and fit the same access pattern as >>>>>> .snapshots or .partitions. The table can stay read-only, limited to >>>>>> the spec-defined shape, and empty when the catalog returns nothing. >>>>>> >>>>>> The tradeoff is that this breaks the current assumption that metadata >>>>>> tables are deterministic projections of table metadata. It’s not >>>>>> explicitly >>>>>> written anywhere, but that assumption exists today. >>>>>> >>>>>> To summarize the questions: >>>>>> >>>>>> 1. >>>>>> >>>>>> Are we OK with catalog-provided metadata tables as a separate >>>>>> category? >>>>>> 2. >>>>>> >>>>>> If yes, should we mark them somehow so they are clearly different >>>>>> from spec-backed metadata tables? >>>>>> 3. >>>>>> >>>>>> If no, what should the SQL surface for labels be instead? >>>>>> >>>>>> Thanks, >>>>>> >>>>>> Andrei >>>>>> >>>>>> [1] REST spec labels: https://github.com/apache/iceberg/pull/15750 >>>>>> [2] SupportsLabels: https://github.com/apache/iceberg/pull/18046 >>>>>> [3] Labels metadata table: >>>>>> https://github.com/apache/iceberg/pull/18048 >>>>>> [4] Spark DESCRIBE: https://github.com/apache/iceberg/pull/18049 >>>>>> [5] Design: >>>>>> https://docs.google.com/document/d/1aj-6JlfBiMYEEVtNuh5WLMOrRQiMCcyYUGbouPM4hXI/edit?tab=t.0#heading=h.2w0kmp1v1gwv >>>>>> >>>>>
