Hi Alex,

Thanks, sounds like multi-version is settled. We both prefer it for the
same upgrade reason, so I’ll keep the remaining discussion to how the
bounds are stored.

I don’t have a strong preference between the current `content_stats` layout
and Russell’s derived `COLLATE(field, collation, version)` approach reusing
`lower_bounds` / `upper_bounds`. In Delta both work the same, so for me
this mostly comes down to on-disk cost.

I’ll prototype both layouts at the physical manifest/stats level and
compare size with 1, 2, and 3 live versions. That should give us something
concrete to discuss at the sync.

I'll set up a dedicated sync for Oct 28 then. Will prepare a small
prototype of layouts by then.

Thanks,
Andrei

On Thu, Oct 1, 2026 at 9:56 PM Alexander Loeser <
[email protected]> wrote:

> Hi Andrei,
>
> Thanks for bringing this up!
>
> I'd also prefer the multi-version approach. Even if engines use only one
> ICU version across the entire fleet, storing a second set of metrics can be
> useful during ICU version upgrades.
>
> I think the exact format we'd use to store the bounds might need further
> discussion. When I talked to Russell a few days back, he seemed worried
> about adding a lot of new fields to content_stats; and instead proposed to
> use derived, non-materialized fields that could have, e.g., a definition of
> COLLATE(field, collation, version), and reuse the lower_bounds and
> upper_bounds fields.
>
> Should we set up another dedicated sync to discuss the best approach?
> Unfortunately I'll be OOO the next two weeks, but how about Oct 28, before
> the community sync like last time?
>
> Best, Alex
>
> On Mon, Sep 21, 2026 at 1:00 PM Andrei Tserakhau via dev <
> [email protected]> wrote:
>
>> Hi all,
>>
>> Splitting one open decision out of the collation thread [1] / PR #16972
>> [2] into its own thread, because it fixes the on-disk metadata layout and
>> I'd like to lock it.
>>
>> Some background: a collated column stores collation-aware min/max bounds
>> so it stays prunable. ICU ordering isn't stable across versions, so each
>> bound is tagged with the collation version it was computed under. A reader
>> uses that bound only if it can reproduce the same version; otherwise it
>> scans the file. That only affects pruning, not correctness.
>>
>> The decision: should one data file be able to carry bounds for more than
>> one collation version?
>>
>>    -
>>
>>    Yes: bounds are a per-file list of {version, lower, upper}. A file
>>    written under ICU X can still be pruned by a reader on ICU Y once a Y 
>> bound
>>    is added, for example by compaction. Cost is a slightly larger stats 
>> struct.
>>    -
>>
>>    No: one version per file. Simpler struct, but a reader on another ICU
>>    version cannot prune that file until it is rewritten. In mixed-version
>>    deployments that can mean a lot more scans during upgrades.
>>
>> My preference is multi-version. In Databricks Runtime customers don't pin
>> an ICU version, so different versions can be live at the same time during
>> upgrades - under one-version-per-file that becomes the common case, not
>> an edge case. That's just the Databricks view, though; I'm asking below
>> because if other engines share the constraint, multi-version is the safer
>> default. The PR already implements it.
>>
>> Two questions, especially for other engines:
>>
>>    1.
>>
>>    In your deployments, is a single known ICU version realistic, or do
>>    you expect several live at once?
>>    2.
>>
>>    Any objection to locking the multi-version layout so we can settle
>>    the field IDs and take the PR out of draft?
>>
>> [1] https://lists.apache.org/thread/44todz4x460g8pb89y8rpozlnmo8vdhc
>> [2] https://github.com/apache/iceberg/pull/16972
>>
>> Thanks,
>> Andrei
>>
>
>
> --
> Alexander Loeser
> Software Engineer
> EMAIL: [email protected]
>
> Snowflake Computing GmbH
> Sitz der Gesellschaft: Leipziger Platz 18, 10117 Berlin, Germany
> <https://www.google.com/maps/search/Leipziger+Platz+18,+10117+Berlin,+Germany?entry=gmail&source=g>
> Geschäftsführer: Erika Lee Payne, Emily Ho
> HRB 241172 B, Amtsgericht Charlottenburg
>
> This message contains information which may be confidential and
> privileged. Unless you are the intended addressee (or authorized to receive
> messages for the intended addressee), you may not use, copy or disclose to
> anyone the message or any information contained in the message or in any
> attachments. If you have received the message in error, please advise the
> sender by reply email and delete the message.
>
>

Reply via email to