Hi Andrei,

Thanks for bringing this up!

I'd also prefer the multi-version approach. Even if engines use only one
ICU version across the entire fleet, storing a second set of metrics can be
useful during ICU version upgrades.

I think the exact format we'd use to store the bounds might need further
discussion. When I talked to Russell a few days back, he seemed worried
about adding a lot of new fields to content_stats; and instead proposed to
use derived, non-materialized fields that could have, e.g., a definition of
COLLATE(field, collation, version), and reuse the lower_bounds and
upper_bounds fields.

Should we set up another dedicated sync to discuss the best approach?
Unfortunately I'll be OOO the next two weeks, but how about Oct 28, before
the community sync like last time?

Best, Alex

On Mon, Sep 21, 2026 at 1:00 PM Andrei Tserakhau via dev <
[email protected]> wrote:

> Hi all,
>
> Splitting one open decision out of the collation thread [1] / PR #16972
> [2] into its own thread, because it fixes the on-disk metadata layout and
> I'd like to lock it.
>
> Some background: a collated column stores collation-aware min/max bounds
> so it stays prunable. ICU ordering isn't stable across versions, so each
> bound is tagged with the collation version it was computed under. A reader
> uses that bound only if it can reproduce the same version; otherwise it
> scans the file. That only affects pruning, not correctness.
>
> The decision: should one data file be able to carry bounds for more than
> one collation version?
>
>    -
>
>    Yes: bounds are a per-file list of {version, lower, upper}. A file
>    written under ICU X can still be pruned by a reader on ICU Y once a Y bound
>    is added, for example by compaction. Cost is a slightly larger stats 
> struct.
>    -
>
>    No: one version per file. Simpler struct, but a reader on another ICU
>    version cannot prune that file until it is rewritten. In mixed-version
>    deployments that can mean a lot more scans during upgrades.
>
> My preference is multi-version. In Databricks Runtime customers don't pin
> an ICU version, so different versions can be live at the same time during
> upgrades - under one-version-per-file that becomes the common case, not
> an edge case. That's just the Databricks view, though; I'm asking below
> because if other engines share the constraint, multi-version is the safer
> default. The PR already implements it.
>
> Two questions, especially for other engines:
>
>    1.
>
>    In your deployments, is a single known ICU version realistic, or do
>    you expect several live at once?
>    2.
>
>    Any objection to locking the multi-version layout so we can settle the
>    field IDs and take the PR out of draft?
>
> [1] https://lists.apache.org/thread/44todz4x460g8pb89y8rpozlnmo8vdhc
> [2] https://github.com/apache/iceberg/pull/16972
>
> Thanks,
> Andrei
>


Alexander Loeser

Reply via email to