Hi Peter,

Thanks for the reply. I think the main question is not about joinability,
but
rather the semantics of metadata tables. The split between table state and
catalog state is the whole point of the design - it's what lets them carry
per-catalog / per-deployment / online context without touching shared state.
Moving them into table state would lose that characteristic.

So I don't think we need to converge on the persistence mechanics here. The
main
question to me: do we want catalog-provided metadata as tables in general?
They
do have a different source and follow different assumptions than the
table-state-backed ones. The question - is this fine?

I think it's a reasonable feature for end-users (see the examples in my
previous
email). We can mitigate the divergent semantics by being explicit about how
these tables are populated and by documenting the contract
(catalog-provided,
point-in-time data, may be empty or vary), so nobody mistakes them for
deterministic metadata.

But I agree it should be a cautious decision.

Best,
Andrei

On Mon, Sep 28, 2026 at 12:41 PM Péter Váry <[email protected]>
wrote:

> Hi Andrei,
>
> Thanks for starting the discussion!
>
> These use-cases push labels towards being fully fledged table metadata:
> queryable, joinable, and expected to be there whenever the table is there.
> That contradicts what we agreed on for the read path, where the spec says
> it "does not require how a catalog produces or stores labels, nor whether
> they are persisted or versioned", and where the goal was to "present the
> current view at the current moment in time".
>
> If we want to go down this road, I think we should reexamine the original
> assumptions, and consider a construct that lives in table metadata, or
> reuse properties for it.
>
> WDYT?
> Peter
>
> Andrei Tserakhau via dev <[email protected]> ezt írta (időpont:
> 2026. szept. 25., P, 23:40):
>
>> Hi Ryan,
>>
>> Agree here: the engine-facing consumption (cost attribution, policy
>> attachment) goes through SupportsLabels, no table needed. The table is for
>> the other consumer - SQL/people or AI Agent :). Some cases that i see here:
>>
>> 1) Exploration. "which columns are classified as X here" is a query a
>> person runs, not something an engine surfaces:
>>
>>     SELECT field_name, key, value
>>     FROM prod.db.orders.labels
>>     WHERE scope = 'field' AND key = 'classification';
>>
>> 2) Composition. .labels joins with .files / .partitions / .snapshots in
>> one query - e.g. attribute bytes to an owner label, i.e. cost attribution
>> as a report someone runs, not a log an engine emits:
>>
>>     SELECT l.value AS owner, SUM(f.file_size_in_bytes) AS bytes
>>     FROM prod.db.orders.files f, prod.db.orders.labels l
>>     WHERE l.scope = 'object' AND l.key = 'owner'
>>     GROUP BY l.value;
>>
>> I think the key value to have a metadata table is joinability, you can’t
>> have it with a programmatic label API.
>>
>> So there is a place for the SQL consumer, the engine path is a different
>> usecase and the co-live together.
>>
>> Thanks,
>> Andrei
>>
>>
>> On Fri, Sep 25, 2026 at 10:51 PM Ryan Blue <[email protected]> wrote:
>>
>>> > Are we OK with catalog-provided metadata tables as a separate category?
>>>
>>> I'm okay with providing metadata through a system table like this, as
>>> long as we think that people will want to access this data that way.
>>> Question 3, "If no, what should the SQL surface for labels be instead?"
>>> makes me think that a SQL surface is _assumed_ to be needed.
>>>
>>> I don't think it is necessarily the case that we need to expose these
>>> for SQL users. I thought that we wanted labels to expose additional context
>>> to engines for things like cost attribution logs or attaching an engine's
>>> policy to a table. That doesn't require a table-like user surface.
>>>
>>> I'm fine adding a metadata table if there's a use for it, but if we
>>> don't need one then it's simpler not to add and maintain it. And that
>>> avoids needing to answer questions like this as well.
>>>
>>>
>>>
>>> On Fri, Sep 25, 2026 at 1:13 PM Andrei Tserakhau via dev <
>>> [email protected]> wrote:
>>>
>>>> Hi all,
>>>>
>>>> Labels in the REST spec recently landed [1]. A catalog can now expose
>>>> object-level and per-field labels on load-table responses. As one of the
>>>> follow-ups, there is a proposal to add a .labels metadata table backed
>>>> by catalog data.
>>>>
>>>> During discussion of the follow-ups ([3], [4]), Peter raised a good
>>>> question on [3]: every metadata table today is derived from table metadata.
>>>> .labels would be different because the data comes from the catalog,
>>>> may vary by catalog, and may be absent if the catalog has nothing to 
>>>> return.
>>>>
>>>> I think this is less about Labels itself and more about a new *kind*
>>>> of data we would expose: metadata owned by the catalog rather than by
>>>> storage. Labels would be the first example, but later the same pattern
>>>> could be used to expose other catalog information as metadata tables.
>>>>
>>>> So the broader question is: do we want metadata tables to also expose
>>>> catalog-provided information?
>>>>
>>>> I think labels are a reasonable first case. They are structured, useful
>>>> to query via SQL, and fit the same access pattern as .snapshots or
>>>> .partitions. The table can stay read-only, limited to the spec-defined
>>>> shape, and empty when the catalog returns nothing.
>>>>
>>>> The tradeoff is that this breaks the current assumption that metadata
>>>> tables are deterministic projections of table metadata. It’s not explicitly
>>>> written anywhere, but that assumption exists today.
>>>>
>>>> To summarize the questions:
>>>>
>>>>    1.
>>>>
>>>>    Are we OK with catalog-provided metadata tables as a separate
>>>>    category?
>>>>    2.
>>>>
>>>>    If yes, should we mark them somehow so they are clearly different
>>>>    from spec-backed metadata tables?
>>>>    3.
>>>>
>>>>    If no, what should the SQL surface for labels be instead?
>>>>
>>>> Thanks,
>>>>
>>>> Andrei
>>>>
>>>> [1] REST spec labels: https://github.com/apache/iceberg/pull/15750
>>>> [2] SupportsLabels: https://github.com/apache/iceberg/pull/18046
>>>> [3] Labels metadata table: https://github.com/apache/iceberg/pull/18048
>>>> [4] Spark DESCRIBE: https://github.com/apache/iceberg/pull/18049
>>>> [5] Design:
>>>> https://docs.google.com/document/d/1aj-6JlfBiMYEEVtNuh5WLMOrRQiMCcyYUGbouPM4hXI/edit?tab=t.0#heading=h.2w0kmp1v1gwv
>>>>
>>>

Reply via email to