Make sure you add it to the dev calendar and send out an announcement email
:)

On Fri, Jul 17, 2026 at 12:42 PM Andrei Tserakhau via dev <
[email protected]> wrote:

> Hi Alexander,
>
> Yes, I think Aug 5th would be ideal, around 5PM CET  / 8AM PST
> I'll make the calendar slot.
>
> Best,
> Andrei
>
> On Fri, Jul 17, 2026 at 3:17 PM Alexander Löser <[email protected]>
> wrote:
>
>> Hi Andrei,
>>
>> thanks for bringing up collations in the last community sync. If I got it
>> right, the next step would be to gather a group of interested folks and set
>> up a dedicated sync - preferably with 2+ weeks headsup so that everyone can
>> plan accordingly.
>>
>> I'm definitely interested in participating in that meeting. How about,
>> e.g., Aug 5th or 7th?
>>
>> Best,
>> Alex
>> On 7/15/26 12:06, Alexander Löser wrote:
>>
>> Hi Andrei,
>>
>> I have some questions/thoughts on the suggestions in your latest mail,
>> but I'm happy to defer those for now in favor of a more general discussion.
>>
>> With regards to your question:
>> > how much cross-engine pruning interoperability should the format
>> guarantee, versus leave to convention?
>>
>> I think this is the right question to ask. In fact, I would set the
>> pruning aspect aside for now and only ask: How much cross-engine
>> interoperability should the format guarantee?
>>
>> If I understand your proposal correctly, you were suggesting that we
>> should allow each engine to choose their ICU version on their own (and
>> skipping pruning if need be), rather than pin a specific ICU version at the
>> table/schema level.
>>
>> I think this suggestion has merit; for example, it will make it much
>> easier for engines to upgrade their ICU version, independently of Iceberg
>> version changes.
>>
>> However, allowing each engine to choose its own ICU version has
>> implications beyond pruning - it also impacts execution.
>> ICU does not guarantee the stability of orderings across different
>> versions, i.e., for two strings x and y, x < y may hold in version N, but
>> not in version N+1. While this usually affects only a small subset of code
>> points, ordering changes have occurred with every other ICU release for the
>> last couple of years.
>> Consequently, the same query may return different results on different
>> engines (using different ICU versions). This can manifest in various ways -
>> different sort orders, more/fewer aggregation results, filtering more/less,
>> etc. - which can be very surprising for users. A colleague mentioned that
>> at a previous company, they had to roll back an ICU library upgrade because
>> their users complained about the sort order differences.
>>
>> Equality deletes are another complication. Until now, whether an equality
>> delete removes a given row is unambiguous - every engine agrees. With
>> different ICU versions, however, engines may draw different conclusions.
>> For example, select * could return different rows if two different ICU
>> versions disagree whether a string matches one of the deleted values. This
>> is, imho, particularly concerning as one DML has the potential to cause
>> different results for all subsequent queries, even if those queries don't
>> use any collated columns. I'm not sure if there would be a way to solve
>> this without disallowing equality deletes on collated columns.
>>
>> Both of the problems described above would disappear if we align on one
>> specified ICU version.
>>
>> To close the loop, I think the question we need to answer is: how much
>> interoperability should the spec guarantee?
>> I don't have a strong opinion, yet - I mostly want to make sure we decide
>> this consciously rather than by omission.
>> I definitely see value in having consistent results across engines.
>> At the same time, I'm not sure if we can expect consistent results across
>> different engines even today: for example, queries involving upper() or
>> lower() may produce different results depending on the case mappings used
>> by the engine - which depend on the Unicode version, too. That said,
>> upper() and lower() are not part of the Iceberg spec while collations would
>> be, so I'm not sure if this is a good reference point.
>>
>> Would be curious to hear what you and the others think.
>>
>> Best, Alex
>>
>>
>> On 7/3/26 15:56, Andrei Tserakhau via dev wrote:
>>
>> Hi Alex,
>>
>> Thanks, these are the right questions. Let me answer them, but I think
>> all three are really facets of one decision worth pulling out, so I'll do
>> that at the end.
>>
>> Original values vs sort keys. I don't think the two limitations are
>> symmetric. You're right that a bound stored under version X may not hold
>> under Y for either representation, and a naive reader prunes only on an
>> exact version match either way. But they degrade differently: a sort key
>> from X is incomparable under Y (nothing a Y reader can do with it) while an
>> original value is the actual string, so a Y reader can re-interpret it. In
>> the case you describe at the end of your mail, where only a small
>> code-point range moved and a file's values fall outside it, that reader can
>> prove the X bound still holds and prune across versions. Sort keys
>> foreclose that; original values keep it open. So original values are a
>> superset: worst case they match sort keys, best case they prune across
>> versions.
>>
>> Your two sort-key advantages are real, I just don't think they belong in
>> the format. Truncation: agreed it's hard for collated strings (your abcเก
>> contraction case is exactly the trap), so I sidestepped it, collation
>> bounds must be tight, a writer that can't store the exact min/max omits the
>> bound. Collation-aware truncation with CollationElementIterator is a
>> possible later optimization. Compare cost: pruning is per-file at planning
>> time, not per-row, so collator vs byte compare is in the noise; and an
>> engine that wants the byte path can derive and cache the sort key from the
>> stored value. Original values don't block that, they just don't bake a
>> version-specific encoding into the format.
>>
>> Column vs file-level version. As you say, every engine can read
>> regardless, so this is a pruning-performance choice, not correctness. In
>> the schema-registered-metrics design a file carries bounds under a declared
>> (collation, version), and a reader prunes any file with a metric for a
>> version it can produce, not only files it wrote. So convergence on one ICU
>> version gives full cross-engine pruning, same as column-level, and a writer
>> or compaction can populate several versions at once. It gives the
>> column-level benefit by convention without making a version bump a
>> format-breaking change.
>>
>> [One data point from the engine side, since I'm coming at this from the
>> Databricks runtime: our runtime is effectively versionless (customers don't
>> pin an ICU version, and upgrades happen under them) so "the same table read
>> by clients on different ICU versions" isn't a corner case for us, it's the
>> default. That's what pushes me toward per-file versioning: pinning one
>> version per table or column means either forcing the whole fleet to upgrade
>> in lockstep or breaking pruning on every bump, and neither survives a
>> versionless fleet. And in practice most version bumps we've gone through
>> don't reorder the data in a given column at all, which is exactly why
>> keeping original values, and eventually your code-point-range idea, lets a
>> reader keep pruning across a bump instead of falling back to a scan.]
>>
>> Providers. Agreed. I'll tighten the spec to a registered set like geo,
>> icu to start, utf8 reserved, non-ICU collations added by spec change rather
>> than ad hoc. Interop is the whole point and an open namespace undercuts it.
>>
>> -----------
>> Stepping back: I think the three above collapse into one question that
>> needs broader alignment than the two of us:
>>
>> > how much cross-engine pruning interoperability should the format
>> guarantee, versus leave to convention?
>>
>> Original-vs-sortkey, column-vs-file version, and open-vs-restricted
>> providers are all that same tradeoff from different angles. It's a values
>> call more than a correctness one, and it binds every engine that hasn't
>> weighed in yet: Trino, Flink, Spark, PyIceberg, rust.
>>
>> From the DBR side I can say the multi-version case is real rather than
>> theoretical, but that's one engine's vantage point. I'd like to get the
>> interop question in front of the other implementers before we fix field
>> IDs, the dev list is probably enough for now, and a community sync is there
>> if it needs more than async. No rush on that; I'd rather let the thread
>> settle the mechanics first.
>>
>> Best,
>> Andrei
>>
>> On Fri, Jul 3, 2026 at 12:22 AM Alexander Löser <[email protected]>
>> wrote:
>>
>>> Hi Andrei,
>>>
>>> Thanks for putting together the spec PR and the detailed write-up! The
>>> approach mostly looks solid to me. I have a few questions/initial thoughts
>>> regarding the changes you proposed (compared to the original proposal):
>>>
>>> > 1 - Bounds store original values, not sort keys, tagged with a
>>> per-file collation version. ICU/CLDR sort keys aren't stable across
>>> versions, so storing keys ties every reader to one exact version; original
>>> values plus a per-file version (readers prune only on an exact match)
>>> degrade gracefully instead of breaking. The schema keeps the collation name
>>> unversioned so anyone can read
>>>
>>> If I understand correctly, we’re talking about two separate things here:
>>>
>>>    1. Tagging a column vs a single file with a certain ICU version
>>>    2. Using collation keys vs original strings (the ones that will
>>>    produce the min/max collation keys)
>>>
>>>
>>> For 1, it comes down to a tradeoff:
>>>
>>>    - If we tag the column with the ICU version, we’d force engines to
>>>    support one agreed-on ICU version if they want to prune files. Engines
>>>    would be able to prune every file (if they support the specific ICU
>>>    version), or none at all, so there is more incentive to support a 
>>> specific
>>>    version
>>>    - If we tag individual files with ICU versions, we gain the big
>>>    advantage that engines do not need to agree on a single ICU version.
>>>    However, if I understand correctly, this comes at the cost of “fractured”
>>>    pruning - engines will only be able to prune files that were written by
>>>    themselves (or rather, with the same ICU version). As a consequence,
>>>    performance might not really be interoperable between different engines.
>>>
>>> Regardless of the approach we choose, all engines should be able to read
>>> the data - they might just not be able to prune files.
>>>
>>> For 2, I’m not sure if I understand the advantages of original strings
>>> yet. As you already pointed out, the collation keys depend on the ICU
>>> version. However, if I understand correctly, the same limitation would
>>> apply to the original strings: the sort order may (and does) change between
>>> different ICU versions, too. As a consequence, we can’t assume the original
>>> lower/upper bound strings we stored for version X will also be lower/upper
>>> bounds for version Y - at least in the general case. So if I understand
>>> correctly, we would not gain additional pruning opportunities compared to
>>> using collation keys. Or am I missing something here?
>>>
>>> At the same time, sort keys do have advantages:
>>>
>>>    - Iceberg allows the truncation of upper- and lower bounds. This is
>>>    trivial for binary collation keys. For original strings, the task becomes
>>>    significantly harder: truncating at a character boundary, for example,
>>>    would lead to wrong results, as there are some context-sensitive 
>>> sequences:
>>>    e.g., with the CLDR root locale, abcเก < abcเ. I think it might be doable
>>>    with ICU’s CollationElementIterator, but it will be tricky to get right.
>>>    - Lower/upper bounds are computed once, but will be compared many
>>>    times. With original strings, we would need to either convert to the
>>>    collation key on the fly, or use ICU’s collator for a direct comparison.
>>>    Both options will be slower than a raw byte-sequence comparison
>>>
>>>
>>> There is one scenario where original values would shine, though. I
>>> analyzed the order-changes between various ICU versions: in many cases,
>>> only a small range of code points changes/is moved. If we had additional
>>> metadata about which code point ranges a file contains (e.g., whether it is
>>> ASCII only), engines might be able to prove that the original string bounds
>>> for version X are still valid for version Y.
>>> If I'm not mistaken, this could allow to prune across different ICU
>>> versions in certain situations, which I’d consider a point in favor of
>>> original values (and file-level ICU versions).
>>>
>>>
>>>
>>> > 2 - A provider-qualified identifier (icu.en_US-ci), leaving room for
>>> non-ICU collations like Spark's UTF8_LCASE, rather than assuming ICU as the
>>> sole provider.
>>>
>>> Adding a provider-mechanism sounds like a good approach to keep the spec
>>> open for future collations [image: :slightly_smiling_face:] I wonder
>>> whether we should restrict the set of allowed providers, though, similar to
>>> how it was done with geo
>>> <https://lists.apache.org/thread/r5x0do8f241bpf565rx8s5s3wc9ogp0f>. My
>>> main motivation for this proposal is interoperability. I worry that
>>> interoperability might suffer or vanish completely if every engine can come
>>> up with their own definitions.
>>>
>>>
>>>
>>> Happy to hear your thoughts on this!
>>>
>>> Best, Alex
>>>
>>>
>>> On 6/27/26 01:37, Szehon Ho wrote:
>>>
>>> Very nice direction, left some comments on the spec proposal.
>>>
>>> Thanks to you folks for working on it !
>>> Szehon
>>>
>>> On Fri, Jun 26, 2026 at 3:29 AM Andrei Tserakhau via dev <
>>> [email protected]> wrote:
>>>
>>>> Hi all,
>>>>
>>>> I've spend some cycle on the collation discussion and make something
>>>> more concrete to react to: a spec-change PR plus reference implementations
>>>> (go and java).
>>>>
>>>> - Spec change (apache/iceberg#16972): a "collation" annotation on
>>>> string fields, and a data_file.collation_bounds field so collated columns
>>>> stay prunable.
>>>> - Reference implementation in iceberg-go (apache/iceberg-go#1318): the
>>>> full path end to end - schema annotation, collation-aware comparison
>>>> (CLDR/UCA), collation bounds in the manifest, and version-gated data-file
>>>> pruning, with an Avro round-trip and pruning tests.
>>>> - A lightweight Java POC (link below): the schema annotation plus a
>>>> Collator-backed comparator, to match where the discussion is. I
>>>> deliberately left the manifest/bounds side out of Java for now.
>>>>
>>>> The design follows the original proposal but takes a few different
>>>> turns, mostly to adopt what we learned in Delta. The ones I'd most like
>>>> input on:
>>>>
>>>> 1 - Bounds store original values, not sort keys, tagged with a per-file
>>>> collation version. ICU/CLDR sort keys aren't stable across versions, so
>>>> storing keys ties every reader to one exact version; original values plus a
>>>> per-file version (readers prune only on an exact match) degrade gracefully
>>>> instead of breaking. The schema keeps the collation name unversioned so
>>>> anyone can read.
>>>>
>>>> 2 - A provider-qualified identifier (icu.en_US-ci), leaving room for
>>>> non-ICU collations like Spark's UTF8_LCASE, rather than assuming ICU as the
>>>> sole provider.
>>>>
>>>> 3 - One structural question I don't have a strong opinion on yet: I put
>>>> collation_bounds on data_file as a standalone v3 field, but field id 146 is
>>>> already the v4 content_stats struct, and collation bounds might belong
>>>> inside that typed-stats framework instead. Worth settling before we fix
>>>> field ids.
>>>>
>>>> The full set of differences and the reader/writer rules are in the PR
>>>> description and the write-up. Comments very welcome — both on the calls
>>>> above and on whether the standalone-field vs content_stats direction is the
>>>> right one.
>>>>
>>>> Best, Andrei
>>>>
>>>> - original proposal:
>>>> https://docs.google.com/document/d/1m8b7u97uteHYjXk-4DNglJSpQO8OcZOCzW2tApCNTW4/edit?tab=t.0
>>>> - spec change: https://github.com/apache/iceberg/pull/16972
>>>> - POC in go: https://github.com/apache/iceberg-go/pull/1318
>>>> - java POC:
>>>> https://github.com/laskoviymishka/iceberg/tree/prototype/collation-support
>>>>
>>>> On Mon, Mar 30, 2026 at 10:54 PM Alexander Löser <
>>>> [email protected]> wrote:
>>>>
>>>>> Hi Andrei,
>>>>>
>>>>> I'm glad you're interested. Looking forward to collaborate with you!
>>>>> Thanks for all the feedback here and in the doc. I only had a quick
>>>>> glance, but I think you raised some good points.  I'll address/respond to
>>>>> your comments as soon as  I get the chance, hopefully tomorrow.
>>>>> I think you also left some comments in this mail that are not yet in
>>>>> the doc - I'll move those to a dedicated section at the end of the doc, so
>>>>> we can use the doc as a single source of truth/discussion.
>>>>>
>>>>> > Happy to share our Delta design doc and implementation learnings in
>>>>> more detail.
>>>>>
>>>>> Sure, sounds good :)
>>>>>
>>>>> Best,
>>>>> Alex
>>>>> On 3/29/26 01:25, Andrei Tserakhau via dev wrote:
>>>>>
>>>>> Hi Alexander,
>>>>>
>>>>> This looks really interesting. We've been working on collation support
>>>>> in Delta and have shipped it in production for some time, so this is an
>>>>> area we care about a lot. If this proposal moves forward we'd be happy to
>>>>> collaborate on the design and implementation.
>>>>>
>>>>> The pseudo-field approach for collation metrics is clean and composes
>>>>> well with existing Iceberg infrastructure. The specifier coverage is
>>>>> comprehensive.
>>>>>
>>>>> A few areas worth discussing as this evolves:
>>>>>
>>>>> 1 - Sort key stability and versioning
>>>>>
>>>>> ICU sort keys are not stable across versions, so a pinned ICU version
>>>>> bump in a future Iceberg release would invalidate all existing collation
>>>>> metrics. In multi-engine environments, requiring all engines to converge 
>>>>> on
>>>>> one ICU version is unrealistic.
>>>>>
>>>>> We store original string values instead of sort keys and allow
>>>>> per-file version annotations -- worth discussing whether something similar
>>>>> could work here.
>>>>>
>>>>> 2 - Provider abstraction
>>>>>
>>>>> The proposal assumes ICU as the sole provider, but Spark ships non-ICU
>>>>> collations like UTF8_LCASE that are widely used. A provider or namespace
>>>>> layer would prevent name collisions and support engine-specific collations
>>>>> without future spec changes.
>>>>>
>>>>> 3 - Operational surface
>>>>>
>>>>> A few things that turned out correctness-critical in our
>>>>> implementation: partition transforms on collated columns (collation-equal
>>>>> but byte-distinct values in different directories), sort order semantics,
>>>>> equality deletes under collation, and Parquet filter pushdown (must be
>>>>> disabled since Parquet has no collation concept).
>>>>>
>>>>> These don't all need to be solved in v1 but would help to scope them.
>>>>>
>>>>> 4 - Smaller items (nit's)
>>>>>
>>>>> UTF-8 bounds for the original field id should be "must write" not
>>>>> "should" -- otherwise backward compat breaks for non-aware engines. Engine
>>>>> fallback behavior (case-sensitive vs older ICU vs fail) could use a
>>>>> recommended preference order to avoid divergent results across engines. 
>>>>> The
>>>>> collation specifier syntax would benefit from a formal grammar.
>>>>>
>>>>> ---
>>>>>
>>>>> Happy to share our Delta design doc and implementation learnings in
>>>>> more detail. Looking forward to the discussion.
>>>>>
>>>>> Best,
>>>>> Andrei
>>>>>
>>>>> On Sat, Mar 28, 2026 at 11:49 PM Alexander Löser <
>>>>> [email protected]> wrote:
>>>>>
>>>>>> Hi everyone,
>>>>>>
>>>>>> this is my first interaction with the Iceberg community, so here a
>>>>>> few words about myself:
>>>>>> - I'm Alex, a Berlin-based software engineer
>>>>>> - I've been working at Snowflake for 4 years now
>>>>>> - I spend most of my time on data types, particularly binary, strings
>>>>>> and collations.
>>>>>>
>>>>>> I'd like to start a discussion about adding collations to the Iceberg
>>>>>> spec.
>>>>>>
>>>>>> Conceptually, collations are an annotation on the string data type.
>>>>>> By default, most engines perform string operations case-sensitively.
>>>>>> Collations allow specifying alternative comparison rules. This is
>>>>>> useful for achieving, e.g., case- or accent-insensitive string 
>>>>>> operations,
>>>>>> or language-specific string sorting.
>>>>>> Collations are supported by many engines: Databricks
>>>>>> <https://docs.databricks.com/aws/en/sql/language-manual/sql-ref-collation>,
>>>>>> Spark
>>>>>> <https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/api/pyspark.sql.functions.collate.html>,
>>>>>> Snowflake <https://docs.snowflake.com/en/sql-reference/collation>,
>>>>>> Oracle
>>>>>> <https://docs.oracle.com/en/database/oracle/oracle-database/19/sqlrf/COLLATION.html>
>>>>>>  - to
>>>>>> name just a few - this list is not complete.
>>>>>>
>>>>>> In Snowflake, we see heavy use of the collation feature. Several
>>>>>> users have approached us, mentioning they want to migrate to Iceberg
>>>>>> tables, but are currently blocked by Iceberg's lack of collation support.
>>>>>>
>>>>>> Given the widespread support for collations across different engines,
>>>>>> I believe introducing collations to Iceberg will increase 
>>>>>> interoperability
>>>>>> and boost its adoption.
>>>>>> I'd be curious about your thoughts.
>>>>>>
>>>>>> *Goal of the proposal*
>>>>>> - Support collation specifications for columns
>>>>>> - Define how collation bounds should be stored - UTF-8 based bounds
>>>>>> are not useful for collated columns
>>>>>>
>>>>>> *Required Changes*
>>>>>> - Extend the schema to let (string) fields be annotated with a
>>>>>> collation
>>>>>>
>>>>>> More details can be found in this doc
>>>>>> <https://docs.google.com/document/d/1m8b7u97uteHYjXk-4DNglJSpQO8OcZOCzW2tApCNTW4/edit?tab=t.0#heading=h.y1ant4w2163k>
>>>>>> .
>>>>>>
>>>>>> I'm also hoping to present the idea in the next community sync.
>>>>>>
>>>>>> Best, Alex
>>>>>>
>>>>>>
>>>>>>

Reply via email to