Hi Manu,

That's a good point, but this is about the labels producer rather than the
table.

The metadata table (as currently implemented) is catalog-agnostic: it reads
the
SupportsLabels capability, not the REST layer directly. So the table is
non-empty for REST catalogs that return labels and empty everywhere else
(including catalogs that don't support labels at all). It's graceful
behaviour,
similar to `.position_deletes` when there are no deletes.

So technically the metadata table works for all catalogs - it's just empty
where
the catalog produces nothing.

Best,
Andrei

On Mon, Sep 28, 2026 at 1:46 PM Manu Zhang <[email protected]> wrote:

> Hi Andrei,
>
> labels landed in REST Spec, but not for all catalogs. On the other hand,
> metadata tables should work for all catalogs, as shown in your example.
> Maybe we need to reframe the main question first.
>
> Regards,
> Manu
>
> Andrei Tserakhau via dev <[email protected]>于2026年9月28日 周一19:32写道:
>
>> Hi Peter,
>>
>> Thanks for the reply. I think the main question is not about joinability,
>> but
>> rather the semantics of metadata tables. The split between table state and
>> catalog state is the whole point of the design - it's what lets them carry
>> per-catalog / per-deployment / online context without touching shared
>> state.
>> Moving them into table state would lose that characteristic.
>>
>> So I don't think we need to converge on the persistence mechanics here.
>> The main
>> question to me: do we want catalog-provided metadata as tables in
>> general? They
>> do have a different source and follow different assumptions than the
>> table-state-backed ones. The question - is this fine?
>>
>> I think it's a reasonable feature for end-users (see the examples in my
>> previous
>> email). We can mitigate the divergent semantics by being explicit about
>> how
>> these tables are populated and by documenting the contract
>> (catalog-provided,
>> point-in-time data, may be empty or vary), so nobody mistakes them for
>> deterministic metadata.
>>
>> But I agree it should be a cautious decision.
>>
>> Best,
>> Andrei
>>
>> On Mon, Sep 28, 2026 at 12:41 PM Péter Váry <[email protected]>
>> wrote:
>>
>>> Hi Andrei,
>>>
>>> Thanks for starting the discussion!
>>>
>>> These use-cases push labels towards being fully fledged table metadata:
>>> queryable, joinable, and expected to be there whenever the table is there.
>>> That contradicts what we agreed on for the read path, where the spec says
>>> it "does not require how a catalog produces or stores labels, nor whether
>>> they are persisted or versioned", and where the goal was to "present the
>>> current view at the current moment in time".
>>>
>>> If we want to go down this road, I think we should reexamine the
>>> original assumptions, and consider a construct that lives in table
>>> metadata, or reuse properties for it.
>>>
>>> WDYT?
>>> Peter
>>>
>>> Andrei Tserakhau via dev <[email protected]> ezt írta (időpont:
>>> 2026. szept. 25., P, 23:40):
>>>
>>>> Hi Ryan,
>>>>
>>>> Agree here: the engine-facing consumption (cost attribution, policy
>>>> attachment) goes through SupportsLabels, no table needed. The table is for
>>>> the other consumer - SQL/people or AI Agent :). Some cases that i see here:
>>>>
>>>> 1) Exploration. "which columns are classified as X here" is a query a
>>>> person runs, not something an engine surfaces:
>>>>
>>>>     SELECT field_name, key, value
>>>>     FROM prod.db.orders.labels
>>>>     WHERE scope = 'field' AND key = 'classification';
>>>>
>>>> 2) Composition. .labels joins with .files / .partitions / .snapshots in
>>>> one query - e.g. attribute bytes to an owner label, i.e. cost attribution
>>>> as a report someone runs, not a log an engine emits:
>>>>
>>>>     SELECT l.value AS owner, SUM(f.file_size_in_bytes) AS bytes
>>>>     FROM prod.db.orders.files f, prod.db.orders.labels l
>>>>     WHERE l.scope = 'object' AND l.key = 'owner'
>>>>     GROUP BY l.value;
>>>>
>>>> I think the key value to have a metadata table is joinability, you
>>>> can’t have it with a programmatic label API.
>>>>
>>>> So there is a place for the SQL consumer, the engine path is a
>>>> different usecase and the co-live together.
>>>>
>>>> Thanks,
>>>> Andrei
>>>>
>>>>
>>>> On Fri, Sep 25, 2026 at 10:51 PM Ryan Blue <[email protected]> wrote:
>>>>
>>>>> > Are we OK with catalog-provided metadata tables as a separate
>>>>> category?
>>>>>
>>>>> I'm okay with providing metadata through a system table like this, as
>>>>> long as we think that people will want to access this data that way.
>>>>> Question 3, "If no, what should the SQL surface for labels be instead?"
>>>>> makes me think that a SQL surface is _assumed_ to be needed.
>>>>>
>>>>> I don't think it is necessarily the case that we need to expose these
>>>>> for SQL users. I thought that we wanted labels to expose additional 
>>>>> context
>>>>> to engines for things like cost attribution logs or attaching an engine's
>>>>> policy to a table. That doesn't require a table-like user surface.
>>>>>
>>>>> I'm fine adding a metadata table if there's a use for it, but if we
>>>>> don't need one then it's simpler not to add and maintain it. And that
>>>>> avoids needing to answer questions like this as well.
>>>>>
>>>>>
>>>>>
>>>>> On Fri, Sep 25, 2026 at 1:13 PM Andrei Tserakhau via dev <
>>>>> [email protected]> wrote:
>>>>>
>>>>>> Hi all,
>>>>>>
>>>>>> Labels in the REST spec recently landed [1]. A catalog can now expose
>>>>>> object-level and per-field labels on load-table responses. As one of the
>>>>>> follow-ups, there is a proposal to add a .labels metadata table
>>>>>> backed by catalog data.
>>>>>>
>>>>>> During discussion of the follow-ups ([3], [4]), Peter raised a good
>>>>>> question on [3]: every metadata table today is derived from table 
>>>>>> metadata.
>>>>>> .labels would be different because the data comes from the catalog,
>>>>>> may vary by catalog, and may be absent if the catalog has nothing to 
>>>>>> return.
>>>>>>
>>>>>> I think this is less about Labels itself and more about a new *kind*
>>>>>> of data we would expose: metadata owned by the catalog rather than by
>>>>>> storage. Labels would be the first example, but later the same pattern
>>>>>> could be used to expose other catalog information as metadata tables.
>>>>>>
>>>>>> So the broader question is: do we want metadata tables to also expose
>>>>>> catalog-provided information?
>>>>>>
>>>>>> I think labels are a reasonable first case. They are structured,
>>>>>> useful to query via SQL, and fit the same access pattern as
>>>>>> .snapshots or .partitions. The table can stay read-only, limited to
>>>>>> the spec-defined shape, and empty when the catalog returns nothing.
>>>>>>
>>>>>> The tradeoff is that this breaks the current assumption that metadata
>>>>>> tables are deterministic projections of table metadata. It’s not 
>>>>>> explicitly
>>>>>> written anywhere, but that assumption exists today.
>>>>>>
>>>>>> To summarize the questions:
>>>>>>
>>>>>>    1.
>>>>>>
>>>>>>    Are we OK with catalog-provided metadata tables as a separate
>>>>>>    category?
>>>>>>    2.
>>>>>>
>>>>>>    If yes, should we mark them somehow so they are clearly different
>>>>>>    from spec-backed metadata tables?
>>>>>>    3.
>>>>>>
>>>>>>    If no, what should the SQL surface for labels be instead?
>>>>>>
>>>>>> Thanks,
>>>>>>
>>>>>> Andrei
>>>>>>
>>>>>> [1] REST spec labels: https://github.com/apache/iceberg/pull/15750
>>>>>> [2] SupportsLabels: https://github.com/apache/iceberg/pull/18046
>>>>>> [3] Labels metadata table:
>>>>>> https://github.com/apache/iceberg/pull/18048
>>>>>> [4] Spark DESCRIBE: https://github.com/apache/iceberg/pull/18049
>>>>>> [5] Design:
>>>>>> https://docs.google.com/document/d/1aj-6JlfBiMYEEVtNuh5WLMOrRQiMCcyYUGbouPM4hXI/edit?tab=t.0#heading=h.2w0kmp1v1gwv
>>>>>>
>>>>>

Reply via email to