Hi all,

we met on Wednesday, and reached an alignment on the direction this proposal should take.

Specifically, these are the directions we agreed on:

 * The ICU version should remain an engine decision; the spec will not
   mandate a specific version. Implications:
     o Writers must tag the bounds they write with the ICU version they
       used
     o Readers must only use bounds written with an ICU version they
       can handle
 * The collation bounds will be represented using original strings
   rather than collation keys
     o on top of the bounds, we will collect an additional metric that
       provides insights over the code points present in a file, which
       can enable engines to perform cross-ICU version pruning
     o the exact shape of that metric still needs to be determined - a
       very simple case could be a flag indicating the file contains
       only ASCII
         + It will be the engine's responsibility to decide whether the
           bounds can be reused across versions. The spec only provides
           information on the code points present in the file
 * Equality deletes will either be deprecated in V4, or not work on
   collated columns

There are some open points that require additional investigation:

 * How should the ICU version be encoded in the bounds? One generated
   expression per collated column per ICU version present might lead to
   a schema explosion
 * How should a code-point related metric look like to enable as much
   cross-ICU version pruning as possible?

I've started updating the original proposal to incorporate the direction we've aligned on. I will also include the tradeoffs we considered and the reasons for our decisions.
I'm planning to share the updated proposal some time next week.

Best,
Alex

On 7/17/26 22:30, Andrei Tserakhau via dev wrote:
Sure!

Scheduled at 5th of August, 5 PM CET / 8 AM PST

Dedicated sync on collation support for the Iceberg spec (PR #16972).

*Goal:* work through the open decisions so we can either commit to a v1 shape or agree it's not worth pursuing.

Decisions to make (in order):

1. *Do we pursue this at all? *Is column-level collation worth adding to the Iceberg spec, given the complexity below, or is it better left to engines? Everything else is moot if the answer is no. 2. *Who owns the ICU version: *format or engine? Pin one version in the format (deterministic across engines, unambiguous equality deletes, but lockstep upgrades) vs. let each engine own its version (independent upgrades, but the same query can return different results on different engines). May split by layer: pruning (performance) vs execution semantics (correctness). 3. *Equality deletes on collated columns:* allow, or disallow in v1? Different ICU versions can disagree whether a delete matched, which can change results for later queries even on non-collated columns. 4. *v1 scope:* pruning + annotation only, or execution semantics too? What's explicitly out (sort orders, partition/bucket transforms)? 5. *Provider model:* restricted registered set like geo (icu to start), or open namespace?

This binds every engine that implements collations, so we want implementers in the room before we fix field IDs. Please come with your engine's ICU-version story: pinned, versionless, and upgrade cadence.

Pre-read:
- Spec PR: iceberg#16972 <https://github.com/apache/iceberg/pull/16972>
- Original proposal: Iceberg Proposal - Collation Support <https://docs.google.com/document/d/1m8b7u97uteHYjXk-4DNglJSpQO8OcZOCzW2tApCNTW4/edit?tab=t.0#heading=h.td5on8bji7s3>

Best,
Andrei

On Fri, Jul 17, 2026 at 8:59 PM Russell Spitzer <[email protected]> wrote:

    Make sure you add it to the dev calendar and send out an
    announcement email :)

    On Fri, Jul 17, 2026 at 12:42 PM Andrei Tserakhau via dev
    <[email protected]> wrote:

        Hi Alexander,

        Yes, I think Aug 5th would be ideal, around 5PM CET  / 8AM PST
        I'll make the calendar slot.

        Best,
        Andrei

        On Fri, Jul 17, 2026 at 3:17 PM Alexander Löser
        <[email protected]> wrote:

            Hi Andrei,

            thanks for bringing up collations in the last community
            sync. If I got it right, the next step would be to gather
            a group of interested folks and set up a dedicated sync -
            preferably with 2+ weeks headsup so that everyone can plan
            accordingly.

            I'm definitely interested in participating in that
            meeting. How about, e.g., Aug 5th or 7th?

            Best,
            Alex

            On 7/15/26 12:06, Alexander Löser wrote:
            Hi Andrei,

            I have some questions/thoughts on the suggestions in your
            latest mail, but I'm happy to defer those for now in
            favor of a more general discussion.

            With regards to your question:
            > how much cross-engine pruning interoperability should
            the format guarantee, versus leave to convention?

            I think this is the right question to ask. In fact, I
            would set the pruning aspect aside for now and only ask:
            How much cross-engine interoperability should the format
            guarantee?

            If I understand your proposal correctly, you were
            suggesting that we should allow each engine to choose
            their ICU version on their own (and skipping pruning if
            need be), rather than pin a specific ICU version at the
            table/schema level.

            I think this suggestion has merit; for example, it will
            make it much easier for engines to upgrade their ICU
            version, independently of Iceberg version changes.

            However, allowing each engine to choose its own ICU
            version has implications beyond pruning - it also impacts
            execution.
            ICU does not guarantee the stability of orderings across
            different versions, i.e., for two strings x and y, x < y
            may hold in version N, but not in version N+1. While this
            usually affects only a small subset of code points,
            ordering changes have occurred with every other ICU
            release for the last couple of years.
            Consequently, the same query may return different results
            on different engines (using different ICU versions). This
            can manifest in various ways - different sort orders,
            more/fewer aggregation results, filtering more/less, etc.
            - which can be very surprising for users. A colleague
            mentioned that at a previous company, they had to roll
            back an ICU library upgrade because their users
            complained about the sort order differences.

            Equality deletes are another complication. Until now,
            whether an equality delete removes a given row is
            unambiguous - every engine agrees. With different ICU
            versions, however, engines may draw different
            conclusions. For example, select * could return different
            rows if two different ICU versions disagree whether a
            string matches one of the deleted values. This is, imho,
            particularly concerning as one DML has the potential to
            cause different results for all subsequent queries, even
            if those queries don't use any collated columns. I'm not
            sure if there would be a way to solve this without
            disallowing equality deletes on collated columns.

            Both of the problems described above would disappear if
            we align on one specified ICU version.

            To close the loop, I think the question we need to answer
            is: how much interoperability should the spec guarantee?
            I don't have a strong opinion, yet - I mostly want to
            make sure we decide this consciously rather than by omission.
            I definitely see value in having consistent results
            across engines.
            At the same time, I'm not sure if we can expect
            consistent results across different engines even today:
            for example, queries involving upper() or lower() may
            produce different results depending on the case mappings
            used by the engine - which depend on the Unicode version,
            too. That said, upper() and lower() are not part of the
            Iceberg spec while collations would be, so I'm not sure
            if this is a good reference point.

            Would be curious to hear what you and the others think.

            Best, Alex


            On 7/3/26 15:56, Andrei Tserakhau via dev wrote:
            Hi Alex,

            Thanks, these are the right questions. Let me answer
            them, but I think all three are really facets of one
            decision worth pulling out, so I'll do that at the end.

            Original values vs sort keys. I don't think the two
            limitations are symmetric. You're right that a bound
            stored under version X may not hold under Y for either
            representation, and a naive reader prunes only on an
            exact version match either way. But they degrade
            differently: a sort key from X is incomparable under Y
            (nothing a Y reader can do with it) while an original
            value is the actual string, so a Y reader can
            re-interpret it. In the case you describe at the end of
            your mail, where only a small code-point range moved and
            a file's values fall outside it, that reader can prove
            the X bound still holds and prune across versions. Sort
            keys foreclose that; original values keep it open. So
            original values are a superset: worst case they match
            sort keys, best case they prune across versions.

            Your two sort-key advantages are real, I just don't
            think they belong in the format. Truncation: agreed it's
            hard for collated strings (your abcเก contraction case
            is exactly the trap), so I sidestepped it, collation
            bounds must be tight, a writer that can't store the
            exact min/max omits the bound. Collation-aware
            truncation with CollationElementIterator is a possible
            later optimization. Compare cost: pruning is per-file at
            planning time, not per-row, so collator vs byte compare
            is in the noise; and an engine that wants the byte path
            can derive and cache the sort key from the stored value.
            Original values don't block that, they just don't bake a
            version-specific encoding into the format.

            Column vs file-level version. As you say, every engine
            can read regardless, so this is a pruning-performance
            choice, not correctness. In the
            schema-registered-metrics design a file carries bounds
            under a declared (collation, version), and a reader
            prunes any file with a metric for a version it can
            produce, not only files it wrote. So convergence on one
            ICU version gives full cross-engine pruning, same as
            column-level, and a writer or compaction can populate
            several versions at once. It gives the column-level
            benefit by convention without making a version bump a
            format-breaking change.

            [One data point from the engine side, since I'm coming
            at this from the Databricks runtime: our runtime is
            effectively versionless (customers don't pin an ICU
            version, and upgrades happen under them) so "the same
            table read by clients on different ICU versions" isn't a
            corner case for us, it's the default. That's what pushes
            me toward per-file versioning: pinning one version per
            table or column means either forcing the whole fleet to
            upgrade in lockstep or breaking pruning on every bump,
            and neither survives a versionless fleet. And in
            practice most version bumps we've gone through don't
            reorder the data in a given column at all, which is
            exactly why keeping original values, and eventually your
            code-point-range idea, lets a reader keep pruning across
            a bump instead of falling back to a scan.]

            Providers. Agreed. I'll tighten the spec to a registered
            set like geo, icu to start, utf8 reserved, non-ICU
            collations added by spec change rather than ad hoc.
            Interop is the whole point and an open namespace
            undercuts it.

            -----------
            Stepping back: I think the three above collapse into one
            question that needs broader alignment than the two of us:

            > how much cross-engine pruning interoperability should
            the format guarantee, versus leave to convention?

            Original-vs-sortkey, column-vs-file version, and
            open-vs-restricted providers are all that same tradeoff
            from different angles. It's a values call more than a
            correctness one, and it binds every engine that hasn't
            weighed in yet: Trino, Flink, Spark, PyIceberg, rust.

            From the DBR side I can say the multi-version case is
            real rather than theoretical, but that's one engine's
            vantage point. I'd like to get the interop question in
            front of the other implementers before we fix field IDs,
            the dev list is probably enough for now, and a community
            sync is there if it needs more than async. No rush on
            that; I'd rather let the thread settle the mechanics first.

            Best,
            Andrei

            On Fri, Jul 3, 2026 at 12:22 AM Alexander Löser
            <[email protected]> wrote:

                Hi Andrei,

                Thanks for putting together the spec PR and the
                detailed write-up! The approach mostly looks solid
                to me. I have a few questions/initial thoughts
                regarding the changes you proposed (compared to the
                original proposal):

                > 1 - Bounds store original values, not sort keys,
                tagged with a per-file collation version. ICU/CLDR
                sort keys aren't stable across versions, so storing
                keys ties every reader to one exact version;
                original values plus a per-file version (readers
                prune only on an exact match) degrade gracefully
                instead of breaking. The schema keeps the collation
                name unversioned so anyone can read

                If I understand correctly, we’re talking about two
                separate things here:

                 1. Tagging a column vs a single file with a certain
                    ICU version
                 2. Using collation keys vs original strings (the
                    ones that will produce the min/max collation keys)


                For 1, it comes down to a tradeoff:

                  * If we tag the column with the ICU version, we’d
                    force engines to support one agreed-on ICU
                    version if they want to prune files. Engines
                    would be able to prune every file (if they
                    support the specific ICU version), or none at
                    all, so there is more incentive to support a
                    specific version
                  * If we tag individual files with ICU versions, we
                    gain the big advantage that engines do not need
                    to agree on a single ICU version. However, if I
                    understand correctly, this comes at the cost of
                    “fractured” pruning - engines will only be able
                    to prune files that were written by themselves
                    (or rather, with the same ICU version). As a
                    consequence, performance might not really be
                    interoperable between different engines.

                Regardless of the approach we choose, all engines
                should be able to read the data - they might just
                not be able to prune files.

                For 2, I’m not sure if I understand the advantages
                of original strings yet. As you already pointed out,
                the collation keys depend on the ICU version.
                However, if I understand correctly, the same
                limitation would apply to the original strings: the
                sort order may (and does) change between different
                ICU versions, too. As a consequence, we can’t assume
                the original lower/upper bound strings we stored for
                version X will also be lower/upper bounds for
                version Y - at least in the general case. So if I
                understand correctly, we would not gain additional
                pruning opportunities compared to using collation
                keys. Or am I missing something here?

                At the same time, sort keys do have advantages:

                  * Iceberg allows the truncation of upper- and
                    lower bounds. This is trivial for binary
                    collation keys. For original strings, the task
                    becomes significantly harder: truncating at a
                    character boundary, for example, would lead to
                    wrong results, as there are some
                    context-sensitive sequences: e.g., with the CLDR
                    root locale, abcเก < abcเ. I think it might be
                    doable with ICU’s CollationElementIterator, but
                    it will be tricky to get right.
                  * Lower/upper bounds are computed once, but will
                    be compared many times. With original strings,
                    we would need to either convert to the collation
                    key on the fly, or use ICU’s collator for a
                    direct comparison. Both options will be slower
                    than a raw byte-sequence comparison


                There is one scenario where original values would
                shine, though. I analyzed the order-changes between
                various ICU versions: in many cases, only a small
                range of code points changes/is moved. If we had
                additional metadata about which code point ranges a
                file contains (e.g., whether it is ASCII only),
                engines might be able to prove that the original
                string bounds for version X are still valid for
                version Y.
                If I'm not mistaken, this could allow to prune
                across different ICU versions in certain situations,
                which I’d consider a point in favor of original
                values (and file-level ICU versions).



                > 2 - A provider-qualified identifier
                (icu.en_US-ci), leaving room for non-ICU collations
                like Spark's UTF8_LCASE, rather than assuming ICU as
                the sole provider.

                Adding a provider-mechanism sounds like a good
                approach to keep the spec open for future collations
                :slightly_smiling_face: I wonder whether we should
                restrict the set of allowed providers, though,
                similar to how it was done with geo
                
<https://lists.apache.org/thread/r5x0do8f241bpf565rx8s5s3wc9ogp0f>.
                My main motivation for this proposal is
                interoperability. I worry that interoperability
                might suffer or vanish completely if every engine
                can come up with their own definitions.



                Happy to hear your thoughts on this!

                Best, Alex


                On 6/27/26 01:37, Szehon Ho wrote:
                Very nice direction, left some comments on the spec
                proposal.

                Thanks to you folks for working on it !
                Szehon

                On Fri, Jun 26, 2026 at 3:29 AM Andrei Tserakhau
                via dev <[email protected]> wrote:

                    Hi all,

                    I've spend some cycle on the collation
                    discussion and make something more concrete to
                    react to: a spec-change PR plus reference
                    implementations (go and java).

                    - Spec change (apache/iceberg#16972): a
                    "collation" annotation on string fields, and a
                    data_file.collation_bounds field so collated
                    columns stay prunable.
                    - Reference implementation in iceberg-go
                    (apache/iceberg-go#1318): the full path end to
                    end - schema annotation, collation-aware
                    comparison (CLDR/UCA), collation bounds in the
                    manifest, and version-gated data-file pruning,
                    with an Avro round-trip and pruning tests.
                    - A lightweight Java POC (link below): the
                    schema annotation plus a Collator-backed
                    comparator, to match where the discussion is. I
                    deliberately left the manifest/bounds side out
                    of Java for now.

                    The design follows the original proposal but
                    takes a few different turns, mostly to adopt
                    what we learned in Delta. The ones I'd most
                    like input on:

                    1 - Bounds store original values, not sort
                    keys, tagged with a per-file collation version.
                    ICU/CLDR sort keys aren't stable across
                    versions, so storing keys ties every reader to
                    one exact version; original values plus a
                    per-file version (readers prune only on an
                    exact match) degrade gracefully instead of
                    breaking. The schema keeps the collation name
                    unversioned so anyone can read.

                    2 - A provider-qualified identifier
                    (icu.en_US-ci), leaving room for non-ICU
                    collations like Spark's UTF8_LCASE, rather than
                    assuming ICU as the sole provider.

                    3 - One structural question I don't have a
                    strong opinion on yet: I put collation_bounds
                    on data_file as a standalone v3 field, but
                    field id 146 is already the v4 content_stats
                    struct, and collation bounds might belong
                    inside that typed-stats framework instead.
                    Worth settling before we fix field ids.

                    The full set of differences and the
                    reader/writer rules are in the PR description
                    and the write-up. Comments very welcome — both
                    on the calls above and on whether the
                    standalone-field vs content_stats direction is
                    the right one.

                    Best, Andrei

                    - original proposal:
                    
https://docs.google.com/document/d/1m8b7u97uteHYjXk-4DNglJSpQO8OcZOCzW2tApCNTW4/edit?tab=t.0
                    - spec change:
                    https://github.com/apache/iceberg/pull/16972
                    - POC in go:
                    https://github.com/apache/iceberg-go/pull/1318
                    - java POC:
                    
https://github.com/laskoviymishka/iceberg/tree/prototype/collation-support

                    On Mon, Mar 30, 2026 at 10:54 PM Alexander
                    Löser <[email protected]> wrote:

                        Hi Andrei,

                        I'm glad you're interested. Looking forward
                        to collaborate with you!
                        Thanks for all the feedback here and in the
                        doc. I only had a quick glance, but I think
                        you raised some good points.  I'll
                        address/respond to your comments as soon
                        as  I get the chance, hopefully tomorrow.
                        I think you also left some comments in this
                        mail that are not yet in the doc - I'll
                        move those to a dedicated section at the
                        end of the doc, so we can use the doc as a
                        single source of truth/discussion.

                        > Happy to share our Delta design doc and
                        implementation learnings in more detail.

                        Sure, sounds good :)

                        Best,
                        Alex

                        On 3/29/26 01:25, Andrei Tserakhau via dev
                        wrote:
                        Hi Alexander,

                        This looks really interesting. We've been
                        working on collation support in Delta and
                        have shipped it in production for some
                        time, so this is an area we care about a
                        lot. If this proposal moves forward we'd
                        be happy to collaborate on the design and
                        implementation.

                        The pseudo-field approach for collation
                        metrics is clean and composes well with
                        existing Iceberg infrastructure. The
                        specifier coverage is comprehensive.

                        A few areas worth discussing as this evolves:

                        1 - Sort key stability and versioning

                        ICU sort keys are not stable across
                        versions, so a pinned ICU version bump in
                        a future Iceberg release would invalidate
                        all existing collation metrics. In
                        multi-engine environments, requiring all
                        engines to converge on one ICU version is
                        unrealistic.

                        We store original string values instead of
                        sort keys and allow per-file version
                        annotations -- worth discussing whether
                        something similar could work here.

                        2 - Provider abstraction

                        The proposal assumes ICU as the sole
                        provider, but Spark ships non-ICU
                        collations like UTF8_LCASE that are widely
                        used. A provider or namespace layer would
                        prevent name collisions and support
                        engine-specific collations without future
                        spec changes.

                        3 - Operational surface

                        A few things that turned out
                        correctness-critical in our
                        implementation: partition transforms on
                        collated columns (collation-equal but
                        byte-distinct values in different
                        directories), sort order semantics,
                        equality deletes under collation, and
                        Parquet filter pushdown (must be disabled
                        since Parquet has no collation concept).

                        These don't all need to be solved in v1
                        but would help to scope them.

                        4 - Smaller items (nit's)

                        UTF-8 bounds for the original field id
                        should be "must write" not "should" --
                        otherwise backward compat breaks for
                        non-aware engines. Engine fallback
                        behavior (case-sensitive vs older ICU vs
                        fail) could use a recommended preference
                        order to avoid divergent results across
                        engines. The collation specifier syntax
                        would benefit from a formal grammar.

                        ---

                        Happy to share our Delta design doc and
                        implementation learnings in more detail.
                        Looking forward to the discussion.

                        Best,
                        Andrei

                        On Sat, Mar 28, 2026 at 11:49 PM Alexander
                        Löser <[email protected]> wrote:

                            Hi everyone,

                            this is my first interaction with the
                            Iceberg community, so here a few words
                            about myself:
                            - I'm Alex, a Berlin-based software
                            engineer
                            - I've been working at Snowflake for 4
                            years now
                            - I spend most of my time on data
                            types, particularly binary, strings
                            and collations.

                            I'd like to start a discussion about
                            adding collations to the Iceberg spec.

                            Conceptually, collations are an
                            annotation on the string data type. By
                            default, most engines perform string
                            operations case-sensitively.
                            Collations allow specifying
                            alternative comparison rules. This is
                            useful for achieving, e.g., case- or
                            accent-insensitive string operations,
                            or language-specific string sorting.
                            Collations are supported by many
                            engines: Databricks
                            
<https://docs.databricks.com/aws/en/sql/language-manual/sql-ref-collation>,
                            Spark
                            
<https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/api/pyspark.sql.functions.collate.html>,
                            Snowflake
                            
<https://docs.snowflake.com/en/sql-reference/collation>,
                            Oracle
                            
<https://docs.oracle.com/en/database/oracle/oracle-database/19/sqlrf/COLLATION.html>
 - to
                            name just a few - this list is not
                            complete.

                            In Snowflake, we see heavy use of the
                            collation feature. Several users have
                            approached us, mentioning they want to
                            migrate to Iceberg tables, but are
                            currently blocked by Iceberg's lack of
                            collation support.

                            Given the widespread support for
                            collations across different engines, I
                            believe introducing collations to
                            Iceberg will increase interoperability
                            and boost its adoption.
                            I'd be curious about your thoughts.

                            *Goal of the proposal*
                            - Support collation specifications for
                            columns
                            - Define how collation bounds should
                            be stored - UTF-8 based bounds are not
                            useful for collated columns

                            *Required Changes*
                            - Extend the schema to let (string)
                            fields be annotated with a collation

                            More details can be found in this doc
                            
<https://docs.google.com/document/d/1m8b7u97uteHYjXk-4DNglJSpQO8OcZOCzW2tApCNTW4/edit?tab=t.0#heading=h.y1ant4w2163k>.

                            I'm also hoping to present the idea in
                            the next community sync.

                            Best, Alex

Reply via email to