Hi Alex, Thanks, sounds like multi-version is settled. We both prefer it for the same upgrade reason, so I’ll keep the remaining discussion to how the bounds are stored.
I don’t have a strong preference between the current `content_stats` layout and Russell’s derived `COLLATE(field, collation, version)` approach reusing `lower_bounds` / `upper_bounds`. In Delta both work the same, so for me this mostly comes down to on-disk cost. I’ll prototype both layouts at the physical manifest/stats level and compare size with 1, 2, and 3 live versions. That should give us something concrete to discuss at the sync. I'll set up a dedicated sync for Oct 28 then. Will prepare a small prototype of layouts by then. Thanks, Andrei On Thu, Oct 1, 2026 at 9:56 PM Alexander Loeser < [email protected]> wrote: > Hi Andrei, > > Thanks for bringing this up! > > I'd also prefer the multi-version approach. Even if engines use only one > ICU version across the entire fleet, storing a second set of metrics can be > useful during ICU version upgrades. > > I think the exact format we'd use to store the bounds might need further > discussion. When I talked to Russell a few days back, he seemed worried > about adding a lot of new fields to content_stats; and instead proposed to > use derived, non-materialized fields that could have, e.g., a definition of > COLLATE(field, collation, version), and reuse the lower_bounds and > upper_bounds fields. > > Should we set up another dedicated sync to discuss the best approach? > Unfortunately I'll be OOO the next two weeks, but how about Oct 28, before > the community sync like last time? > > Best, Alex > > On Mon, Sep 21, 2026 at 1:00 PM Andrei Tserakhau via dev < > [email protected]> wrote: > >> Hi all, >> >> Splitting one open decision out of the collation thread [1] / PR #16972 >> [2] into its own thread, because it fixes the on-disk metadata layout and >> I'd like to lock it. >> >> Some background: a collated column stores collation-aware min/max bounds >> so it stays prunable. ICU ordering isn't stable across versions, so each >> bound is tagged with the collation version it was computed under. A reader >> uses that bound only if it can reproduce the same version; otherwise it >> scans the file. That only affects pruning, not correctness. >> >> The decision: should one data file be able to carry bounds for more than >> one collation version? >> >> - >> >> Yes: bounds are a per-file list of {version, lower, upper}. A file >> written under ICU X can still be pruned by a reader on ICU Y once a Y >> bound >> is added, for example by compaction. Cost is a slightly larger stats >> struct. >> - >> >> No: one version per file. Simpler struct, but a reader on another ICU >> version cannot prune that file until it is rewritten. In mixed-version >> deployments that can mean a lot more scans during upgrades. >> >> My preference is multi-version. In Databricks Runtime customers don't pin >> an ICU version, so different versions can be live at the same time during >> upgrades - under one-version-per-file that becomes the common case, not >> an edge case. That's just the Databricks view, though; I'm asking below >> because if other engines share the constraint, multi-version is the safer >> default. The PR already implements it. >> >> Two questions, especially for other engines: >> >> 1. >> >> In your deployments, is a single known ICU version realistic, or do >> you expect several live at once? >> 2. >> >> Any objection to locking the multi-version layout so we can settle >> the field IDs and take the PR out of draft? >> >> [1] https://lists.apache.org/thread/44todz4x460g8pb89y8rpozlnmo8vdhc >> [2] https://github.com/apache/iceberg/pull/16972 >> >> Thanks, >> Andrei >> > > > -- > Alexander Loeser > Software Engineer > EMAIL: [email protected] > > Snowflake Computing GmbH > Sitz der Gesellschaft: Leipziger Platz 18, 10117 Berlin, Germany > <https://www.google.com/maps/search/Leipziger+Platz+18,+10117+Berlin,+Germany?entry=gmail&source=g> > Geschäftsführer: Erika Lee Payne, Emily Ho > HRB 241172 B, Amtsgericht Charlottenburg > > This message contains information which may be confidential and > privileged. Unless you are the intended addressee (or authorized to receive > messages for the intended addressee), you may not use, copy or disclose to > anyone the message or any information contained in the message or in any > attachments. If you have received the message in error, please advise the > sender by reply email and delete the message. > >
