Index would also use a "Sort Order" functionality.
If you take a look at the proposed spec, I have defined Cluster spec
<https://github.com/apache/iceberg/pull/16961/changes#diff-327b0c238ffecd1fd8a6c9a87cdc9bc9c5d552203851997acdaf5c1dc0c1f795R143>
as
a list of materialized, and non-materialized fields, and sort the data in
the index based on that. If we redefine sort order, then we can reuse that
with the index too.
We could also reuse the Group max value
<https://github.com/apache/iceberg/pull/16961/changes#diff-327b0c238ffecd1fd8a6c9a87cdc9bc9c5d552203851997acdaf5c1dc0c1f795R322>
and
maybe add Group min value statistics too, which would define the ranges
available in the given data file.

This proposal would make expressing these things way easier!

Thanks,
Peter

Gianluca Graziadei <[email protected]> ezt írta (időpont: 2026.
szept. 3., Cs, 22:13):

> +1 (non-binding)
> Thanks, that's exactly the distinction I was trying to get at. If the
> parametrization is fully captured by the expression in metadata, then we're
> aligned: there are no hidden inputs and the expression remains a
> well-defined transformation of the row.
>
> - Gianluca
>
> Il giorno gio 3 set 2026 alle ore 21:58 Ryan Blue <[email protected]> ha
> scritto:
>
>> Thanks for the background and hilbert details. I understand what you're
>> saying, but I don't think it affects this work. However we parameterize
>> this function, what we store in metadata needs to completely describe the
>> function's operation. We are concerned with functions that are well-defined
>> transformations of a single input row. That could be something like
>> `hilbert(zvalue(col1, 0, 1), zvalue(col2, 100, 50))` to account for what
>> you're talking about.
>>
>> On Thu, Sep 3, 2026 at 12:17 PM Gianluca Graziadei <
>> [email protected]> wrote:
>>
>>> Hi Ryan,
>>>
>>> You're right, and apologies: I left the actual case implicit and used
>>> hilbert(c1, c2) as shorthand for it. With a fixed per-column bit width the
>>> curve index is a pure function of the row, so as written my example doesn't
>>> support my point. Let me state the real case explicitly.
>>>
>>> Hilbert clustering over heterogeneous columns doesn't work on raw
>>> values. Each column has a different domain, scale and distribution, and if
>>> you feed them to the curve as-is, most of the bits of one dimension end up
>>> unused and the clustering degenerates along it.
>>> This is not hypothetical: it came out of the discussion with Russell on
>>> the Hilbert PR, and it's what the PR 17893 implements on top of the
>>> MultiColumn Term refactor (see also this draft
>>> https://github.com/GGraziadei/iceberg/pull/1).
>>>
>>> So in practice each input goes through a per-column transform that maps
>>> it onto the available bits, and the expression is: *hilbert(linear(c1),
>>> zstdnorm(c2), quantile(c3)) *where linear is a min/max rescaling,
>>> zstdnorm a z-score standardization, and quantile a mapping onto quantile
>>> boundaries.
>>>
>>> I don't think this needs a new structure because the proposal already
>>> fits it, provided two things are explicit in the spec (RFC 2119 ):
>>>
>>> - expr MUST be a deterministic function of the referenced field values
>>> only; any constant, including fitted parameters, MUST be embedded as a
>>> literal in the expression. This is what makes "no hidden inputs" a
>>> checkable property rather than an assumption.
>>> - Entries in the non-materialized fields table MUST be immutable:
>>> changing an expression, including re-fitting its parameters, allocates a
>>> new field ID, and IDs MUST NOT be reused (maybe this is implicit).
>>>
>>> I think this is also where Russell's two points land. A sort order over
>>> an expression like the one above is precisely the case where you want to
>>> save the computed values, and the lineage question is the transformation
>>> re-fit case: under rule 2 the field ID itself carries the lineage, so I
>>> don't think a separate ExpressionsID is needed. If instead an entry can be
>>> updated in place under the same ID, then we do need versioning. Is
>>> immutability the intent?
>>>
>>> I'll park the CDF suggestion (I will detach in a separate thread that
>>> could be useful also for optimizing the Spark physical plan): it's about
>>> stats representation, not id allocation however it is important at this
>>> level to understand this exigence (=validity scope for expressions) and
>>> plan according to it.
>>>
>>> Cheers,
>>> Gianluca
>>>
>>> Il giorno gio 3 set 2026 alle ore 20:23 Russell Spitzer <
>>> [email protected]> ha scritto:
>>>
>>>> This makes sense for me. I think as we do this we should also define an
>>>> output for "Sort Orders" since that's the other main function I think we
>>>> have that we would want to save values for and it should look pretty
>>>> similar to partition output (just mimicking the input values).)
>>>>
>>>>
>>>> Since these wouldn't live in the "Schema' are we also tracking the
>>>> lineage of these expressions over time? Does a table have both a SchemaID
>>>> and ExpressionsID?
>>>>
>>>> On Thu, Sep 3, 2026 at 1:17 PM Ryan Blue <[email protected]> wrote:
>>>>
>>>>> Thanks for taking a look, Peter and Gianluca.
>>>>>
>>>>> For the result data type, I was thinking that we would require it if
>>>>> it is not a case where we know it can be derived. For instance, we know 
>>>>> the
>>>>> output type for partition fields so we don't need to specify it and it can
>>>>> change. For cases where you may not know the output, like when you have an
>>>>> expression, we would require the data type. And I agree that we would
>>>>> use bound or id-based references.
>>>>>
>>>>> I don't think that there is much value in having a special case for
>>>>> identity. You already have a field ID for the data, so I don't think there
>>>>> is a situation in which you would ever do this.
>>>>>
>>>>> Also, I see the distinction that Gianluca points out, but this is not
>>>>> an issue because there should be no functions in the second category. 
>>>>> Maybe
>>>>> hilbert was a poor choice for an example. I don't think that we need to
>>>>> design for expressions that have hidden inputs.
>>>>>
>>>>> Ryan
>>>>>
>>>>> On Tue, Sep 1, 2026 at 4:17 AM Péter Váry <[email protected]>
>>>>> wrote:
>>>>>
>>>>>> Hi Ryan,
>>>>>>
>>>>>> Thanks for putting this together. This would make index definitions
>>>>>> significantly simpler.
>>>>>>
>>>>>> A couple of thoughts:
>>>>>>
>>>>>>    - *Schema evolution*: In Iceberg expressions, the result type is
>>>>>>    currently defined by the expression itself and may evolve as the table
>>>>>>    schema changes (e.g., int -> long). Storing an explicit data-type 
>>>>>> would
>>>>>>    make reading metadata and index data much simpler, but it does remove 
>>>>>> some
>>>>>>    of that flexibility.
>>>>>>    - *FieldId references*: It would probably be a good idea to
>>>>>>    require expressions to reference fields by ID (e.g., field(15)) 
>>>>>> rather than
>>>>>>    by name (col_a) to make them resilient to column renames.
>>>>>>    - *Identity expressions*: I expect index definitions to
>>>>>>    frequently contain entries like:
>>>>>>
>>>>>>
>>>>>> *              {"field-id": 104, "type": "expr-value", "data-type":
>>>>>>    "long", "expr": "<identity(col_a)>"} *
>>>>>>    Would it be worth introducing a dedicated type for this case,
>>>>>>    such as:
>>>>>>
>>>>>>
>>>>>> *             {"field-id": 104, "type": "identity-value",
>>>>>>    "original-field-id": 15} *
>>>>>>    It doesn't seem broadly useful outside indexing, but identity
>>>>>>    projections in indexes are likely common enough that a more compact
>>>>>>    representation may be worth considering.
>>>>>>
>>>>>> Thanks,
>>>>>> Peter
>>>>>>
>>>>>> Gianluca Graziadei <[email protected]> ezt írta (időpont:
>>>>>> 2026. szept. 1., K, 7:02):
>>>>>>
>>>>>>> Hi Ryan,
>>>>>>> I like the proposal; it is concise and clear.
>>>>>>>
>>>>>>> I would suggest making a clear distinction between two classes of
>>>>>>> expressions:
>>>>>>>
>>>>>>> 1. Expressions that produce deterministic values intrinsically tied
>>>>>>> to the input value, e.g. to_lower_case(s).This class is relatively
>>>>>>> straightforward to reason about, since the result depends only on the
>>>>>>> individual row.
>>>>>>> 2. Expressions such as Hilbert that produce values whose meaning
>>>>>>> depends on the input value + the whole distribution. Here the main 
>>>>>>> issue is
>>>>>>> that the underlying distribution matters. If the distribution changes 
>>>>>>> over
>>>>>>> time, hilbert(c11,c12) computed at time t may no longer be valid at t+1.
>>>>>>> This makes the second class harder to handle, because we need to 
>>>>>>> establish
>>>>>>> for how long a computed value remains valid (=if in the current snapshot
>>>>>>> values computed on a previous snapshot are still valid).
>>>>>>>
>>>>>>> For this second class, I don't think persistence of the computed
>>>>>>> value alone is sufficient. We may also need to persist the distribution 
>>>>>>> (or
>>>>>>> its CDF) against which the value was computed (I am not confident that
>>>>>>> min/max can be good enough in this case).
>>>>>>>
>>>>>>> Before I comment:
>>>>>>> 1. Are you targeting to manage these two classes in the same way? It
>>>>>>> seems that they have potentially different validity scope.
>>>>>>> 2. What about adding a snapshot level CDF struct per column?
>>>>>>>
>>>>>>> Cheers,
>>>>>>> Gianluca
>>>>>>>
>>>>>>> Il giorno mar 1 set 2026 alle ore 00:25 Ryan Blue <[email protected]>
>>>>>>> ha scritto:
>>>>>>>
>>>>>>>> Hi everyone,
>>>>>>>>
>>>>>>>> One of the remaining open questions for v4 metadata is how we will
>>>>>>>> assign table field IDs for values that are not written into the table. 
>>>>>>>> I
>>>>>>>> want to propose a solution that I think is going to be flexible, while 
>>>>>>>> not
>>>>>>>> introducing a lot of churn in the table or REST specs.
>>>>>>>>
>>>>>>>> Columnar field stats are written into metadata using a simple
>>>>>>>> function from table field ID to metadata field ID. We want to reuse 
>>>>>>>> what we
>>>>>>>> already have working for table fields and keep the spec simple. That 
>>>>>>>> means
>>>>>>>> we need a way to assign a table field ID to a non-materialized column 
>>>>>>>> so
>>>>>>>> that we can track its stats. There are a few cases we’ve identified:
>>>>>>>>
>>>>>>>>    - Partition field output for non-monotonic functions, like 
>>>>>>>> bucket(1024,
>>>>>>>>    id)
>>>>>>>>    - Clustering expressions, like to_lower_case(last_name)
>>>>>>>>    - Collation sequence lower and upper bounds
>>>>>>>>
>>>>>>>> We also discussed a new case this morning in the index sync: we
>>>>>>>> need a field ID for a derived value used to organize an index, like 
>>>>>>>> hilbert(col_a,
>>>>>>>> col_b), because we intend to use table field IDs in index schemas.
>>>>>>>>
>>>>>>>> Initially, I suggested that we keep a table of expressions and
>>>>>>>> assign each one a field ID. But as we started thinking about the use 
>>>>>>>> cases
>>>>>>>> where we need expressions it became clear that denormalizing *all*
>>>>>>>> expressions was adding a lot of complexity for little benefit. For 
>>>>>>>> example,
>>>>>>>> CHECK constraints won’t have reusable expressions and it makes
>>>>>>>> little sense to create them in two parts (expression and constraint).
>>>>>>>> Similarly, it is awkward to model a collation sequence as an 
>>>>>>>> expression,
>>>>>>>> and we don’t need to rebuild partition specs just to assign field IDs.
>>>>>>>> However, we also don’t want to just embed table field IDs in every one 
>>>>>>>> of
>>>>>>>> these structures.
>>>>>>>>
>>>>>>>> My proposal is to directly model what we want: one table of fields
>>>>>>>> that are not materialized in the table, but are assigned IDs for stats 
>>>>>>>> or
>>>>>>>> other purposes. This would take a few forms:
>>>>>>>>
>>>>>>>>    - Partition output value: {"field-id": 102, "type":
>>>>>>>>    "partition-value", "partition-field-id": 1000}
>>>>>>>>    - Collation sequence: {"field-id": 103, "type":
>>>>>>>>    "collation-bounds", "collation-seq-id": 1}
>>>>>>>>    - Value expression results: {"field-id": 104, "type":
>>>>>>>>    "expr-value", "data-type": "long", "expr": <hilbert(col_a, col_b) 
>>>>>>>> expr>}
>>>>>>>>
>>>>>>>> This representation leaves existing structures alone and is a
>>>>>>>> single place outside of schema to allocate table field IDs. This can be
>>>>>>>> expanded with new types later when we want to add new structures, like 
>>>>>>>> a
>>>>>>>> cluster-by spec.
>>>>>>>>
>>>>>>>> I think this is a fairly clean way to move forward and solve two
>>>>>>>> challenges that we’re currently hitting. We'll discuss this in the 
>>>>>>>> next v4
>>>>>>>> metadata sync, but in the meantime please reply with feedback if you 
>>>>>>>> have
>>>>>>>> an opinion.
>>>>>>>>
>>>>>>>> Thanks,
>>>>>>>>
>>>>>>>> Ryan
>>>>>>>>
>>>>>>>

Reply via email to