Re: [Discussion] Collation Support

2026-07-17 Thread Andrei Tserakhau via dev
Sure!

Scheduled at 5th of August, 5 PM CET / 8 AM PST

Dedicated sync on collation support for the Iceberg spec (PR #16972).

*Goal:* work through the open decisions so we can either commit to a v1
shape or agree it's not worth pursuing.

Decisions to make (in order):

1. *Do we pursue this at all? *Is column-level collation worth adding to
the Iceberg spec, given the complexity below, or is it better left to
engines? Everything else is moot if the answer is no.
2. *Who owns the ICU version: *format or engine? Pin one version in the
format (deterministic across engines, unambiguous equality deletes, but
lockstep upgrades) vs. let each engine own its version (independent
upgrades, but the same query can return different results on different
engines). May split by layer: pruning (performance) vs execution semantics
(correctness).
3. *Equality deletes on collated columns:* allow, or disallow in v1?
Different ICU versions can disagree whether a delete matched, which can
change results for later queries even on non-collated columns.
4. *v1 scope:* pruning + annotation only, or execution semantics too?
What's explicitly out (sort orders, partition/bucket transforms)?
5. *Provider model:* restricted registered set like geo (icu to start), or
open namespace?

This binds every engine that implements collations, so we want implementers
in the room before we fix field IDs. Please come with your engine's
ICU-version story: pinned, versionless, and upgrade cadence.

Pre-read:
- Spec PR: iceberg#16972 
- Original proposal: Iceberg Proposal - Collation Support


Best,
Andrei

On Fri, Jul 17, 2026 at 8:59 PM Russell Spitzer 
wrote:

> Make sure you add it to the dev calendar and send out an announcement
> email :)
>
> On Fri, Jul 17, 2026 at 12:42 PM Andrei Tserakhau via dev <
> [email protected]> wrote:
>
>> Hi Alexander,
>>
>> Yes, I think Aug 5th would be ideal, around 5PM CET  / 8AM PST
>> I'll make the calendar slot.
>>
>> Best,
>> Andrei
>>
>> On Fri, Jul 17, 2026 at 3:17 PM Alexander Löser 
>> wrote:
>>
>>> Hi Andrei,
>>>
>>> thanks for bringing up collations in the last community sync. If I got
>>> it right, the next step would be to gather a group of interested folks and
>>> set up a dedicated sync - preferably with 2+ weeks headsup so that everyone
>>> can plan accordingly.
>>>
>>> I'm definitely interested in participating in that meeting. How about,
>>> e.g., Aug 5th or 7th?
>>>
>>> Best,
>>> Alex
>>> On 7/15/26 12:06, Alexander Löser wrote:
>>>
>>> Hi Andrei,
>>>
>>> I have some questions/thoughts on the suggestions in your latest mail,
>>> but I'm happy to defer those for now in favor of a more general discussion.
>>>
>>> With regards to your question:
>>> > how much cross-engine pruning interoperability should the format
>>> guarantee, versus leave to convention?
>>>
>>> I think this is the right question to ask. In fact, I would set the
>>> pruning aspect aside for now and only ask: How much cross-engine
>>> interoperability should the format guarantee?
>>>
>>> If I understand your proposal correctly, you were suggesting that we
>>> should allow each engine to choose their ICU version on their own (and
>>> skipping pruning if need be), rather than pin a specific ICU version at the
>>> table/schema level.
>>>
>>> I think this suggestion has merit; for example, it will make it much
>>> easier for engines to upgrade their ICU version, independently of Iceberg
>>> version changes.
>>>
>>> However, allowing each engine to choose its own ICU version has
>>> implications beyond pruning - it also impacts execution.
>>> ICU does not guarantee the stability of orderings across different
>>> versions, i.e., for two strings x and y, x < y may hold in version N, but
>>> not in version N+1. While this usually affects only a small subset of code
>>> points, ordering changes have occurred with every other ICU release for the
>>> last couple of years.
>>> Consequently, the same query may return different results on different
>>> engines (using different ICU versions). This can manifest in various ways -
>>> different sort orders, more/fewer aggregation results, filtering more/less,
>>> etc. - which can be very surprising for users. A colleague mentioned that
>>> at a previous company, they had to roll back an ICU library upgrade because
>>> their users complained about the sort order differences.
>>>
>>> Equality deletes are another complication. Until now, whether an
>>> equality delete removes a given row is unambiguous - every engine agrees.
>>> With different ICU versions, however, engines may draw different
>>> conclusions. For example, select * could return different rows if two
>>> different ICU versions disagree whether a string matches one of the deleted
>>> values. This is, imho, particularly concerning as one DML has the potential
>>> to cause di

Re: [Discussion] Collation Support

2026-07-17 Thread Russell Spitzer
Make sure you add it to the dev calendar and send out an announcement email
:)

On Fri, Jul 17, 2026 at 12:42 PM Andrei Tserakhau via dev <
[email protected]> wrote:

> Hi Alexander,
>
> Yes, I think Aug 5th would be ideal, around 5PM CET  / 8AM PST
> I'll make the calendar slot.
>
> Best,
> Andrei
>
> On Fri, Jul 17, 2026 at 3:17 PM Alexander Löser 
> wrote:
>
>> Hi Andrei,
>>
>> thanks for bringing up collations in the last community sync. If I got it
>> right, the next step would be to gather a group of interested folks and set
>> up a dedicated sync - preferably with 2+ weeks headsup so that everyone can
>> plan accordingly.
>>
>> I'm definitely interested in participating in that meeting. How about,
>> e.g., Aug 5th or 7th?
>>
>> Best,
>> Alex
>> On 7/15/26 12:06, Alexander Löser wrote:
>>
>> Hi Andrei,
>>
>> I have some questions/thoughts on the suggestions in your latest mail,
>> but I'm happy to defer those for now in favor of a more general discussion.
>>
>> With regards to your question:
>> > how much cross-engine pruning interoperability should the format
>> guarantee, versus leave to convention?
>>
>> I think this is the right question to ask. In fact, I would set the
>> pruning aspect aside for now and only ask: How much cross-engine
>> interoperability should the format guarantee?
>>
>> If I understand your proposal correctly, you were suggesting that we
>> should allow each engine to choose their ICU version on their own (and
>> skipping pruning if need be), rather than pin a specific ICU version at the
>> table/schema level.
>>
>> I think this suggestion has merit; for example, it will make it much
>> easier for engines to upgrade their ICU version, independently of Iceberg
>> version changes.
>>
>> However, allowing each engine to choose its own ICU version has
>> implications beyond pruning - it also impacts execution.
>> ICU does not guarantee the stability of orderings across different
>> versions, i.e., for two strings x and y, x < y may hold in version N, but
>> not in version N+1. While this usually affects only a small subset of code
>> points, ordering changes have occurred with every other ICU release for the
>> last couple of years.
>> Consequently, the same query may return different results on different
>> engines (using different ICU versions). This can manifest in various ways -
>> different sort orders, more/fewer aggregation results, filtering more/less,
>> etc. - which can be very surprising for users. A colleague mentioned that
>> at a previous company, they had to roll back an ICU library upgrade because
>> their users complained about the sort order differences.
>>
>> Equality deletes are another complication. Until now, whether an equality
>> delete removes a given row is unambiguous - every engine agrees. With
>> different ICU versions, however, engines may draw different conclusions.
>> For example, select * could return different rows if two different ICU
>> versions disagree whether a string matches one of the deleted values. This
>> is, imho, particularly concerning as one DML has the potential to cause
>> different results for all subsequent queries, even if those queries don't
>> use any collated columns. I'm not sure if there would be a way to solve
>> this without disallowing equality deletes on collated columns.
>>
>> Both of the problems described above would disappear if we align on one
>> specified ICU version.
>>
>> To close the loop, I think the question we need to answer is: how much
>> interoperability should the spec guarantee?
>> I don't have a strong opinion, yet - I mostly want to make sure we decide
>> this consciously rather than by omission.
>> I definitely see value in having consistent results across engines.
>> At the same time, I'm not sure if we can expect consistent results across
>> different engines even today: for example, queries involving upper() or
>> lower() may produce different results depending on the case mappings used
>> by the engine - which depend on the Unicode version, too. That said,
>> upper() and lower() are not part of the Iceberg spec while collations would
>> be, so I'm not sure if this is a good reference point.
>>
>> Would be curious to hear what you and the others think.
>>
>> Best, Alex
>>
>>
>> On 7/3/26 15:56, Andrei Tserakhau via dev wrote:
>>
>> Hi Alex,
>>
>> Thanks, these are the right questions. Let me answer them, but I think
>> all three are really facets of one decision worth pulling out, so I'll do
>> that at the end.
>>
>> Original values vs sort keys. I don't think the two limitations are
>> symmetric. You're right that a bound stored under version X may not hold
>> under Y for either representation, and a naive reader prunes only on an
>> exact version match either way. But they degrade differently: a sort key
>> from X is incomparable under Y (nothing a Y reader can do with it) while an
>> original value is the actual string, so a Y reader can re-interpret it. In
>> the case you describe at the

Re: [Discussion] Collation Support

2026-07-17 Thread Andrei Tserakhau via dev
Hi Alexander,

Yes, I think Aug 5th would be ideal, around 5PM CET  / 8AM PST
I'll make the calendar slot.

Best,
Andrei

On Fri, Jul 17, 2026 at 3:17 PM Alexander Löser 
wrote:

> Hi Andrei,
>
> thanks for bringing up collations in the last community sync. If I got it
> right, the next step would be to gather a group of interested folks and set
> up a dedicated sync - preferably with 2+ weeks headsup so that everyone can
> plan accordingly.
>
> I'm definitely interested in participating in that meeting. How about,
> e.g., Aug 5th or 7th?
>
> Best,
> Alex
> On 7/15/26 12:06, Alexander Löser wrote:
>
> Hi Andrei,
>
> I have some questions/thoughts on the suggestions in your latest mail, but
> I'm happy to defer those for now in favor of a more general discussion.
>
> With regards to your question:
> > how much cross-engine pruning interoperability should the format
> guarantee, versus leave to convention?
>
> I think this is the right question to ask. In fact, I would set the
> pruning aspect aside for now and only ask: How much cross-engine
> interoperability should the format guarantee?
>
> If I understand your proposal correctly, you were suggesting that we
> should allow each engine to choose their ICU version on their own (and
> skipping pruning if need be), rather than pin a specific ICU version at the
> table/schema level.
>
> I think this suggestion has merit; for example, it will make it much
> easier for engines to upgrade their ICU version, independently of Iceberg
> version changes.
>
> However, allowing each engine to choose its own ICU version has
> implications beyond pruning - it also impacts execution.
> ICU does not guarantee the stability of orderings across different
> versions, i.e., for two strings x and y, x < y may hold in version N, but
> not in version N+1. While this usually affects only a small subset of code
> points, ordering changes have occurred with every other ICU release for the
> last couple of years.
> Consequently, the same query may return different results on different
> engines (using different ICU versions). This can manifest in various ways -
> different sort orders, more/fewer aggregation results, filtering more/less,
> etc. - which can be very surprising for users. A colleague mentioned that
> at a previous company, they had to roll back an ICU library upgrade because
> their users complained about the sort order differences.
>
> Equality deletes are another complication. Until now, whether an equality
> delete removes a given row is unambiguous - every engine agrees. With
> different ICU versions, however, engines may draw different conclusions.
> For example, select * could return different rows if two different ICU
> versions disagree whether a string matches one of the deleted values. This
> is, imho, particularly concerning as one DML has the potential to cause
> different results for all subsequent queries, even if those queries don't
> use any collated columns. I'm not sure if there would be a way to solve
> this without disallowing equality deletes on collated columns.
>
> Both of the problems described above would disappear if we align on one
> specified ICU version.
>
> To close the loop, I think the question we need to answer is: how much
> interoperability should the spec guarantee?
> I don't have a strong opinion, yet - I mostly want to make sure we decide
> this consciously rather than by omission.
> I definitely see value in having consistent results across engines.
> At the same time, I'm not sure if we can expect consistent results across
> different engines even today: for example, queries involving upper() or
> lower() may produce different results depending on the case mappings used
> by the engine - which depend on the Unicode version, too. That said,
> upper() and lower() are not part of the Iceberg spec while collations would
> be, so I'm not sure if this is a good reference point.
>
> Would be curious to hear what you and the others think.
>
> Best, Alex
>
>
> On 7/3/26 15:56, Andrei Tserakhau via dev wrote:
>
> Hi Alex,
>
> Thanks, these are the right questions. Let me answer them, but I think all
> three are really facets of one decision worth pulling out, so I'll do that
> at the end.
>
> Original values vs sort keys. I don't think the two limitations are
> symmetric. You're right that a bound stored under version X may not hold
> under Y for either representation, and a naive reader prunes only on an
> exact version match either way. But they degrade differently: a sort key
> from X is incomparable under Y (nothing a Y reader can do with it) while an
> original value is the actual string, so a Y reader can re-interpret it. In
> the case you describe at the end of your mail, where only a small
> code-point range moved and a file's values fall outside it, that reader can
> prove the X bound still holds and prune across versions. Sort keys
> foreclose that; original values keep it open. So original values are a
> superset: worst case they

Re: [Discussion] Collation Support

2026-07-17 Thread Alexander Löser

Hi Andrei,

thanks for bringing up collations in the last community sync. If I got 
it right, the next step would be to gather a group of interested folks 
and set up a dedicated sync - preferably with 2+ weeks headsup so that 
everyone can plan accordingly.


I'm definitely interested in participating in that meeting. How about, 
e.g., Aug 5th or 7th?


Best,
Alex

On 7/15/26 12:06, Alexander Löser wrote:

Hi Andrei,

I have some questions/thoughts on the suggestions in your latest mail, 
but I'm happy to defer those for now in favor of a more general 
discussion.


With regards to your question:
> how much cross-engine pruning interoperability should the format 
guarantee, versus leave to convention?


I think this is the right question to ask. In fact, I would set the 
pruning aspect aside for now and only ask: How much cross-engine 
interoperability should the format guarantee?


If I understand your proposal correctly, you were suggesting that we 
should allow each engine to choose their ICU version on their own (and 
skipping pruning if need be), rather than pin a specific ICU version 
at the table/schema level.


I think this suggestion has merit; for example, it will make it much 
easier for engines to upgrade their ICU version, independently of 
Iceberg version changes.


However, allowing each engine to choose its own ICU version has 
implications beyond pruning - it also impacts execution.
ICU does not guarantee the stability of orderings across different 
versions, i.e., for two strings x and y, x < y may hold in version N, 
but not in version N+1. While this usually affects only a small subset 
of code points, ordering changes have occurred with every other ICU 
release for the last couple of years.
Consequently, the same query may return different results on different 
engines (using different ICU versions). This can manifest in various 
ways - different sort orders, more/fewer aggregation results, 
filtering more/less, etc. - which can be very surprising for users. A 
colleague mentioned that at a previous company, they had to roll back 
an ICU library upgrade because their users complained about the sort 
order differences.


Equality deletes are another complication. Until now, whether an 
equality delete removes a given row is unambiguous - every engine 
agrees. With different ICU versions, however, engines may draw 
different conclusions. For example, select * could return different 
rows if two different ICU versions disagree whether a string matches 
one of the deleted values. This is, imho, particularly concerning as 
one DML has the potential to cause different results for all 
subsequent queries, even if those queries don't use any collated 
columns. I'm not sure if there would be a way to solve this without 
disallowing equality deletes on collated columns.


Both of the problems described above would disappear if we align on 
one specified ICU version.


To close the loop, I think the question we need to answer is: how much 
interoperability should the spec guarantee?
I don't have a strong opinion, yet - I mostly want to make sure we 
decide this consciously rather than by omission.

I definitely see value in having consistent results across engines.
At the same time, I'm not sure if we can expect consistent results 
across different engines even today: for example, queries involving 
upper() or lower() may produce different results depending on the case 
mappings used by the engine - which depend on the Unicode version, 
too. That said, upper() and lower() are not part of the Iceberg spec 
while collations would be, so I'm not sure if this is a good reference 
point.


Would be curious to hear what you and the others think.

Best, Alex


On 7/3/26 15:56, Andrei Tserakhau via dev wrote:

Hi Alex,

Thanks, these are the right questions. Let me answer them, but I 
think all three are really facets of one decision worth pulling out, 
so I'll do that at the end.


Original values vs sort keys. I don't think the two limitations are 
symmetric. You're right that a bound stored under version X may not 
hold under Y for either representation, and a naive reader prunes 
only on an exact version match either way. But they degrade 
differently: a sort key from X is incomparable under Y (nothing a Y 
reader can do with it) while an original value is the actual string, 
so a Y reader can re-interpret it. In the case you describe at the 
end of your mail, where only a small code-point range moved and a 
file's values fall outside it, that reader can prove the X bound 
still holds and prune across versions. Sort keys foreclose that; 
original values keep it open. So original values are a superset: 
worst case they match sort keys, best case they prune across versions.


Your two sort-key advantages are real, I just don't think they belong 
in the format. Truncation: agreed it's hard for collated strings 
(your abcเก contraction case is exactly the trap), so I sidestepped 
it, collation bounds m

Re: [Discussion] Collation Support

2026-07-15 Thread Alexander Löser

Hi Andrei,

I have some questions/thoughts on the suggestions in your latest mail, 
but I'm happy to defer those for now in favor of a more general discussion.


With regards to your question:
> how much cross-engine pruning interoperability should the format 
guarantee, versus leave to convention?


I think this is the right question to ask. In fact, I would set the 
pruning aspect aside for now and only ask: How much cross-engine 
interoperability should the format guarantee?


If I understand your proposal correctly, you were suggesting that we 
should allow each engine to choose their ICU version on their own (and 
skipping pruning if need be), rather than pin a specific ICU version at 
the table/schema level.


I think this suggestion has merit; for example, it will make it much 
easier for engines to upgrade their ICU version, independently of 
Iceberg version changes.


However, allowing each engine to choose its own ICU version has 
implications beyond pruning - it also impacts execution.
ICU does not guarantee the stability of orderings across different 
versions, i.e., for two strings x and y, x < y may hold in version N, 
but not in version N+1. While this usually affects only a small subset 
of code points, ordering changes have occurred with every other ICU 
release for the last couple of years.
Consequently, the same query may return different results on different 
engines (using different ICU versions). This can manifest in various 
ways - different sort orders, more/fewer aggregation results, filtering 
more/less, etc. - which can be very surprising for users. A colleague 
mentioned that at a previous company, they had to roll back an ICU 
library upgrade because their users complained about the sort order 
differences.


Equality deletes are another complication. Until now, whether an 
equality delete removes a given row is unambiguous - every engine 
agrees. With different ICU versions, however, engines may draw different 
conclusions. For example, select * could return different rows if two 
different ICU versions disagree whether a string matches one of the 
deleted values. This is, imho, particularly concerning as one DML has 
the potential to cause different results for all subsequent queries, 
even if those queries don't use any collated columns. I'm not sure if 
there would be a way to solve this without disallowing equality deletes 
on collated columns.


Both of the problems described above would disappear if we align on one 
specified ICU version.


To close the loop, I think the question we need to answer is: how much 
interoperability should the spec guarantee?
I don't have a strong opinion, yet - I mostly want to make sure we 
decide this consciously rather than by omission.

I definitely see value in having consistent results across engines.
At the same time, I'm not sure if we can expect consistent results 
across different engines even today: for example, queries involving 
upper() or lower() may produce different results depending on the case 
mappings used by the engine - which depend on the Unicode version, too. 
That said, upper() and lower() are not part of the Iceberg spec while 
collations would be, so I'm not sure if this is a good reference point.


Would be curious to hear what you and the others think.

Best, Alex


On 7/3/26 15:56, Andrei Tserakhau via dev wrote:

Hi Alex,

Thanks, these are the right questions. Let me answer them, but I think 
all three are really facets of one decision worth pulling out, so I'll 
do that at the end.


Original values vs sort keys. I don't think the two limitations are 
symmetric. You're right that a bound stored under version X may not 
hold under Y for either representation, and a naive reader prunes only 
on an exact version match either way. But they degrade differently: a 
sort key from X is incomparable under Y (nothing a Y reader can do 
with it) while an original value is the actual string, so a Y reader 
can re-interpret it. In the case you describe at the end of your mail, 
where only a small code-point range moved and a file's values fall 
outside it, that reader can prove the X bound still holds and prune 
across versions. Sort keys foreclose that; original values keep it 
open. So original values are a superset: worst case they match sort 
keys, best case they prune across versions.


Your two sort-key advantages are real, I just don't think they belong 
in the format. Truncation: agreed it's hard for collated strings (your 
abcเก contraction case is exactly the trap), so I sidestepped it, 
collation bounds must be tight, a writer that can't store the exact 
min/max omits the bound. Collation-aware truncation with 
CollationElementIterator is a possible later optimization. Compare 
cost: pruning is per-file at planning time, not per-row, so collator 
vs byte compare is in the noise; and an engine that wants the byte 
path can derive and cache the sort key from the stored value. Original 
values don't block that, t

Re: [Discussion] Collation Support

2026-07-03 Thread Andrei Tserakhau via dev
Hi Alex,

Thanks, these are the right questions. Let me answer them, but I think all
three are really facets of one decision worth pulling out, so I'll do that
at the end.

Original values vs sort keys. I don't think the two limitations are
symmetric. You're right that a bound stored under version X may not hold
under Y for either representation, and a naive reader prunes only on an
exact version match either way. But they degrade differently: a sort key
from X is incomparable under Y (nothing a Y reader can do with it) while an
original value is the actual string, so a Y reader can re-interpret it. In
the case you describe at the end of your mail, where only a small
code-point range moved and a file's values fall outside it, that reader can
prove the X bound still holds and prune across versions. Sort keys
foreclose that; original values keep it open. So original values are a
superset: worst case they match sort keys, best case they prune across
versions.

Your two sort-key advantages are real, I just don't think they belong in
the format. Truncation: agreed it's hard for collated strings (your abcเก
contraction case is exactly the trap), so I sidestepped it, collation
bounds must be tight, a writer that can't store the exact min/max omits the
bound. Collation-aware truncation with CollationElementIterator is a
possible later optimization. Compare cost: pruning is per-file at planning
time, not per-row, so collator vs byte compare is in the noise; and an
engine that wants the byte path can derive and cache the sort key from the
stored value. Original values don't block that, they just don't bake a
version-specific encoding into the format.

Column vs file-level version. As you say, every engine can read regardless,
so this is a pruning-performance choice, not correctness. In the
schema-registered-metrics design a file carries bounds under a declared
(collation, version), and a reader prunes any file with a metric for a
version it can produce, not only files it wrote. So convergence on one ICU
version gives full cross-engine pruning, same as column-level, and a writer
or compaction can populate several versions at once. It gives the
column-level benefit by convention without making a version bump a
format-breaking change.

[One data point from the engine side, since I'm coming at this from the
Databricks runtime: our runtime is effectively versionless (customers don't
pin an ICU version, and upgrades happen under them) so "the same table read
by clients on different ICU versions" isn't a corner case for us, it's the
default. That's what pushes me toward per-file versioning: pinning one
version per table or column means either forcing the whole fleet to upgrade
in lockstep or breaking pruning on every bump, and neither survives a
versionless fleet. And in practice most version bumps we've gone through
don't reorder the data in a given column at all, which is exactly why
keeping original values, and eventually your code-point-range idea, lets a
reader keep pruning across a bump instead of falling back to a scan.]

Providers. Agreed. I'll tighten the spec to a registered set like geo, icu
to start, utf8 reserved, non-ICU collations added by spec change rather
than ad hoc. Interop is the whole point and an open namespace undercuts it.

---
Stepping back: I think the three above collapse into one question that
needs broader alignment than the two of us:

> how much cross-engine pruning interoperability should the format
guarantee, versus leave to convention?

Original-vs-sortkey, column-vs-file version, and open-vs-restricted
providers are all that same tradeoff from different angles. It's a values
call more than a correctness one, and it binds every engine that hasn't
weighed in yet: Trino, Flink, Spark, PyIceberg, rust.

>From the DBR side I can say the multi-version case is real rather than
theoretical, but that's one engine's vantage point. I'd like to get the
interop question in front of the other implementers before we fix field
IDs, the dev list is probably enough for now, and a community sync is there
if it needs more than async. No rush on that; I'd rather let the thread
settle the mechanics first.

Best,
Andrei

On Fri, Jul 3, 2026 at 12:22 AM Alexander Löser 
wrote:

> Hi Andrei,
>
> Thanks for putting together the spec PR and the detailed write-up! The
> approach mostly looks solid to me. I have a few questions/initial thoughts
> regarding the changes you proposed (compared to the original proposal):
>
> > 1 - Bounds store original values, not sort keys, tagged with a per-file
> collation version. ICU/CLDR sort keys aren't stable across versions, so
> storing keys ties every reader to one exact version; original values plus a
> per-file version (readers prune only on an exact match) degrade gracefully
> instead of breaking. The schema keeps the collation name unversioned so
> anyone can read
>
> If I understand correctly, we’re talking about two separate things here:
>
>1. Tagging a colu

Re: [Discussion] Collation Support

2026-07-02 Thread Alexander Löser

Hi Andrei,

Thanks for putting together the spec PR and the detailed write-up! The 
approach mostly looks solid to me. I have a few questions/initial 
thoughts regarding the changes you proposed (compared to the original 
proposal):


> 1 - Bounds store original values, not sort keys, tagged with a 
per-file collation version. ICU/CLDR sort keys aren't stable across 
versions, so storing keys ties every reader to one exact version; 
original values plus a per-file version (readers prune only on an exact 
match) degrade gracefully instead of breaking. The schema keeps the 
collation name unversioned so anyone can read


If I understand correctly, we’re talking about two separate things here:

1. Tagging a column vs a single file with a certain ICU version
2. Using collation keys vs original strings (the ones that will produce
   the min/max collation keys)


For 1, it comes down to a tradeoff:

 * If we tag the column with the ICU version, we’d force engines to
   support one agreed-on ICU version if they want to prune files.
   Engines would be able to prune every file (if they support the
   specific ICU version), or none at all, so there is more incentive to
   support a specific version
 * If we tag individual files with ICU versions, we gain the big
   advantage that engines do not need to agree on a single ICU version.
   However, if I understand correctly, this comes at the cost of
   “fractured” pruning - engines will only be able to prune files that
   were written by themselves (or rather, with the same ICU version).
   As a consequence, performance might not really be interoperable
   between different engines.

Regardless of the approach we choose, all engines should be able to read 
the data - they might just not be able to prune files.


For 2, I’m not sure if I understand the advantages of original strings 
yet. As you already pointed out, the collation keys depend on the ICU 
version. However, if I understand correctly, the same limitation would 
apply to the original strings: the sort order may (and does) change 
between different ICU versions, too. As a consequence, we can’t assume 
the original lower/upper bound strings we stored for version X will also 
be lower/upper bounds for version Y - at least in the general case. So 
if I understand correctly, we would not gain additional pruning 
opportunities compared to using collation keys. Or am I missing 
something here?


At the same time, sort keys do have advantages:

 * Iceberg allows the truncation of upper- and lower bounds. This is
   trivial for binary collation keys. For original strings, the task
   becomes significantly harder: truncating at a character boundary,
   for example, would lead to wrong results, as there are some
   context-sensitive sequences: e.g., with the CLDR root locale, abcเก
   < abcเ. I think it might be doable with ICU’s
   CollationElementIterator, but it will be tricky to get right.
 * Lower/upper bounds are computed once, but will be compared many
   times. With original strings, we would need to either convert to the
   collation key on the fly, or use ICU’s collator for a direct
   comparison. Both options will be slower than a raw byte-sequence
   comparison


There is one scenario where original values would shine, though. I 
analyzed the order-changes between various ICU versions: in many cases, 
only a small range of code points changes/is moved. If we had additional 
metadata about which code point ranges a file contains (e.g., whether it 
is ASCII only), engines might be able to prove that the original string 
bounds for version X are still valid for version Y.
If I'm not mistaken, this could allow to prune across different ICU 
versions in certain situations, which I’d consider a point in favor of 
original values (and file-level ICU versions).




> 2 - A provider-qualified identifier (icu.en_US-ci), leaving room for 
non-ICU collations like Spark's UTF8_LCASE, rather than assuming ICU as 
the sole provider.


Adding a provider-mechanism sounds like a good approach to keep the spec 
open for future collations :slightly_smiling_face: I wonder whether we 
should restrict the set of allowed providers, though, similar to how it 
was done with geo 
. My 
main motivation for this proposal is interoperability. I worry that 
interoperability might suffer or vanish completely if every engine can 
come up with their own definitions.




Happy to hear your thoughts on this!

Best, Alex


On 6/27/26 01:37, Szehon Ho wrote:

Very nice direction, left some comments on the spec proposal.

Thanks to you folks for working on it !
Szehon

On Fri, Jun 26, 2026 at 3:29 AM Andrei Tserakhau via dev 
 wrote:


Hi all,

I've spend some cycle on the collation discussion and make
something more concrete to react to: a spec-change PR plus
reference implementations (go and java).

- Spec change (apache/iceberg#16972): a "collation" annota

Re: [Discussion] Collation Support

2026-06-26 Thread Szehon Ho
Very nice direction, left some comments on the spec proposal.

Thanks to you folks for working on it !
Szehon

On Fri, Jun 26, 2026 at 3:29 AM Andrei Tserakhau via dev <
[email protected]> wrote:

> Hi all,
>
> I've spend some cycle on the collation discussion and make something more
> concrete to react to: a spec-change PR plus reference implementations (go
> and java).
>
> - Spec change (apache/iceberg#16972): a "collation" annotation on string
> fields, and a data_file.collation_bounds field so collated columns stay
> prunable.
> - Reference implementation in iceberg-go (apache/iceberg-go#1318): the
> full path end to end - schema annotation, collation-aware comparison
> (CLDR/UCA), collation bounds in the manifest, and version-gated data-file
> pruning, with an Avro round-trip and pruning tests.
> - A lightweight Java POC (link below): the schema annotation plus a
> Collator-backed comparator, to match where the discussion is. I
> deliberately left the manifest/bounds side out of Java for now.
>
> The design follows the original proposal but takes a few different turns,
> mostly to adopt what we learned in Delta. The ones I'd most like input on:
>
> 1 - Bounds store original values, not sort keys, tagged with a per-file
> collation version. ICU/CLDR sort keys aren't stable across versions, so
> storing keys ties every reader to one exact version; original values plus a
> per-file version (readers prune only on an exact match) degrade gracefully
> instead of breaking. The schema keeps the collation name unversioned so
> anyone can read.
>
> 2 - A provider-qualified identifier (icu.en_US-ci), leaving room for
> non-ICU collations like Spark's UTF8_LCASE, rather than assuming ICU as the
> sole provider.
>
> 3 - One structural question I don't have a strong opinion on yet: I put
> collation_bounds on data_file as a standalone v3 field, but field id 146 is
> already the v4 content_stats struct, and collation bounds might belong
> inside that typed-stats framework instead. Worth settling before we fix
> field ids.
>
> The full set of differences and the reader/writer rules are in the PR
> description and the write-up. Comments very welcome — both on the calls
> above and on whether the standalone-field vs content_stats direction is the
> right one.
>
> Best, Andrei
>
> - original proposal:
> https://docs.google.com/document/d/1m8b7u97uteHYjXk-4DNglJSpQO8OcZOCzW2tApCNTW4/edit?tab=t.0
> - spec change: https://github.com/apache/iceberg/pull/16972
> - POC in go: https://github.com/apache/iceberg-go/pull/1318
> - java POC:
> https://github.com/laskoviymishka/iceberg/tree/prototype/collation-support
>
> On Mon, Mar 30, 2026 at 10:54 PM Alexander Löser 
> wrote:
>
>> Hi Andrei,
>>
>> I'm glad you're interested. Looking forward to collaborate with you!
>> Thanks for all the feedback here and in the doc. I only had a quick
>> glance, but I think you raised some good points.  I'll address/respond to
>> your comments as soon as  I get the chance, hopefully tomorrow.
>> I think you also left some comments in this mail that are not yet in the
>> doc - I'll move those to a dedicated section at the end of the doc, so we
>> can use the doc as a single source of truth/discussion.
>>
>> > Happy to share our Delta design doc and implementation learnings in
>> more detail.
>>
>> Sure, sounds good :)
>>
>> Best,
>> Alex
>> On 3/29/26 01:25, Andrei Tserakhau via dev wrote:
>>
>> Hi Alexander,
>>
>> This looks really interesting. We've been working on collation support in
>> Delta and have shipped it in production for some time, so this is an area
>> we care about a lot. If this proposal moves forward we'd be happy to
>> collaborate on the design and implementation.
>>
>> The pseudo-field approach for collation metrics is clean and composes
>> well with existing Iceberg infrastructure. The specifier coverage is
>> comprehensive.
>>
>> A few areas worth discussing as this evolves:
>>
>> 1 - Sort key stability and versioning
>>
>> ICU sort keys are not stable across versions, so a pinned ICU version
>> bump in a future Iceberg release would invalidate all existing collation
>> metrics. In multi-engine environments, requiring all engines to converge on
>> one ICU version is unrealistic.
>>
>> We store original string values instead of sort keys and allow per-file
>> version annotations -- worth discussing whether something similar could
>> work here.
>>
>> 2 - Provider abstraction
>>
>> The proposal assumes ICU as the sole provider, but Spark ships non-ICU
>> collations like UTF8_LCASE that are widely used. A provider or namespace
>> layer would prevent name collisions and support engine-specific collations
>> without future spec changes.
>>
>> 3 - Operational surface
>>
>> A few things that turned out correctness-critical in our implementation:
>> partition transforms on collated columns (collation-equal but byte-distinct
>> values in different directories), sort order semantics, equality deletes
>> under collation, and Parquet

Re: [Discussion] Collation Support

2026-06-26 Thread Andrei Tserakhau via dev
Hi all,

I've spend some cycle on the collation discussion and make something more
concrete to react to: a spec-change PR plus reference implementations (go
and java).

- Spec change (apache/iceberg#16972): a "collation" annotation on string
fields, and a data_file.collation_bounds field so collated columns stay
prunable.
- Reference implementation in iceberg-go (apache/iceberg-go#1318): the full
path end to end - schema annotation, collation-aware comparison (CLDR/UCA),
collation bounds in the manifest, and version-gated data-file pruning, with
an Avro round-trip and pruning tests.
- A lightweight Java POC (link below): the schema annotation plus a
Collator-backed comparator, to match where the discussion is. I
deliberately left the manifest/bounds side out of Java for now.

The design follows the original proposal but takes a few different turns,
mostly to adopt what we learned in Delta. The ones I'd most like input on:

1 - Bounds store original values, not sort keys, tagged with a per-file
collation version. ICU/CLDR sort keys aren't stable across versions, so
storing keys ties every reader to one exact version; original values plus a
per-file version (readers prune only on an exact match) degrade gracefully
instead of breaking. The schema keeps the collation name unversioned so
anyone can read.

2 - A provider-qualified identifier (icu.en_US-ci), leaving room for
non-ICU collations like Spark's UTF8_LCASE, rather than assuming ICU as the
sole provider.

3 - One structural question I don't have a strong opinion on yet: I put
collation_bounds on data_file as a standalone v3 field, but field id 146 is
already the v4 content_stats struct, and collation bounds might belong
inside that typed-stats framework instead. Worth settling before we fix
field ids.

The full set of differences and the reader/writer rules are in the PR
description and the write-up. Comments very welcome — both on the calls
above and on whether the standalone-field vs content_stats direction is the
right one.

Best, Andrei

- original proposal:
https://docs.google.com/document/d/1m8b7u97uteHYjXk-4DNglJSpQO8OcZOCzW2tApCNTW4/edit?tab=t.0
- spec change: https://github.com/apache/iceberg/pull/16972
- POC in go: https://github.com/apache/iceberg-go/pull/1318
- java POC:
https://github.com/laskoviymishka/iceberg/tree/prototype/collation-support

On Mon, Mar 30, 2026 at 10:54 PM Alexander Löser 
wrote:

> Hi Andrei,
>
> I'm glad you're interested. Looking forward to collaborate with you!
> Thanks for all the feedback here and in the doc. I only had a quick
> glance, but I think you raised some good points.  I'll address/respond to
> your comments as soon as  I get the chance, hopefully tomorrow.
> I think you also left some comments in this mail that are not yet in the
> doc - I'll move those to a dedicated section at the end of the doc, so we
> can use the doc as a single source of truth/discussion.
>
> > Happy to share our Delta design doc and implementation learnings in more
> detail.
>
> Sure, sounds good :)
>
> Best,
> Alex
> On 3/29/26 01:25, Andrei Tserakhau via dev wrote:
>
> Hi Alexander,
>
> This looks really interesting. We've been working on collation support in
> Delta and have shipped it in production for some time, so this is an area
> we care about a lot. If this proposal moves forward we'd be happy to
> collaborate on the design and implementation.
>
> The pseudo-field approach for collation metrics is clean and composes well
> with existing Iceberg infrastructure. The specifier coverage is
> comprehensive.
>
> A few areas worth discussing as this evolves:
>
> 1 - Sort key stability and versioning
>
> ICU sort keys are not stable across versions, so a pinned ICU version bump
> in a future Iceberg release would invalidate all existing collation
> metrics. In multi-engine environments, requiring all engines to converge on
> one ICU version is unrealistic.
>
> We store original string values instead of sort keys and allow per-file
> version annotations -- worth discussing whether something similar could
> work here.
>
> 2 - Provider abstraction
>
> The proposal assumes ICU as the sole provider, but Spark ships non-ICU
> collations like UTF8_LCASE that are widely used. A provider or namespace
> layer would prevent name collisions and support engine-specific collations
> without future spec changes.
>
> 3 - Operational surface
>
> A few things that turned out correctness-critical in our implementation:
> partition transforms on collated columns (collation-equal but byte-distinct
> values in different directories), sort order semantics, equality deletes
> under collation, and Parquet filter pushdown (must be disabled since
> Parquet has no collation concept).
>
> These don't all need to be solved in v1 but would help to scope them.
>
> 4 - Smaller items (nit's)
>
> UTF-8 bounds for the original field id should be "must write" not "should"
> -- otherwise backward compat breaks for non-aware engines. Engine fallback
> behavior (case-

Re: [Discussion] Collation Support

2026-03-30 Thread Alexander Löser

Hi Andrei,

I'm glad you're interested. Looking forward to collaborate with you!
Thanks for all the feedback here and in the doc. I only had a quick 
glance, but I think you raised some good points.  I'll address/respond 
to your comments as soon as  I get the chance, hopefully tomorrow.
I think you also left some comments in this mail that are not yet in the 
doc - I'll move those to a dedicated section at the end of the doc, so 
we can use the doc as a single source of truth/discussion.


> Happy to share our Delta design doc and implementation learnings in 
more detail.


Sure, sounds good :)

Best,
Alex

On 3/29/26 01:25, Andrei Tserakhau via dev wrote:

Hi Alexander,

This looks really interesting. We've been working on collation support 
in Delta and have shipped it in production for some time, so this is 
an area we care about a lot. If this proposal moves forward we'd be 
happy to collaborate on the design and implementation.


The pseudo-field approach for collation metrics is clean and composes 
well with existing Iceberg infrastructure. The specifier coverage is 
comprehensive.


A few areas worth discussing as this evolves:

1 - Sort key stability and versioning

ICU sort keys are not stable across versions, so a pinned ICU version 
bump in a future Iceberg release would invalidate all existing 
collation metrics. In multi-engine environments, requiring all engines 
to converge on one ICU version is unrealistic.


We store original string values instead of sort keys and allow 
per-file version annotations -- worth discussing whether something 
similar could work here.


2 - Provider abstraction

The proposal assumes ICU as the sole provider, but Spark ships non-ICU 
collations like UTF8_LCASE that are widely used. A provider or 
namespace layer would prevent name collisions and support 
engine-specific collations without future spec changes.


3 - Operational surface

A few things that turned out correctness-critical in our 
implementation: partition transforms on collated columns 
(collation-equal but byte-distinct values in different directories), 
sort order semantics, equality deletes under collation, and Parquet 
filter pushdown (must be disabled since Parquet has no collation 
concept).


These don't all need to be solved in v1 but would help to scope them.

4 - Smaller items (nit's)

UTF-8 bounds for the original field id should be "must write" not 
"should" -- otherwise backward compat breaks for non-aware engines. 
Engine fallback behavior (case-sensitive vs older ICU vs fail) could 
use a recommended preference order to avoid divergent results across 
engines. The collation specifier syntax would benefit from a formal 
grammar.


---

Happy to share our Delta design doc and implementation learnings in 
more detail. Looking forward to the discussion.


Best,
Andrei

On Sat, Mar 28, 2026 at 11:49 PM Alexander Löser 
 wrote:


Hi everyone,

this is my first interaction with the Iceberg community, so here a
few words about myself:
- I'm Alex, a Berlin-based software engineer
- I've been working at Snowflake for 4 years now
- I spend most of my time on data types, particularly binary,
strings and collations.

I'd like to start a discussion about adding collations to the
Iceberg spec.

Conceptually, collations are an annotation on the string data
type. By default, most engines perform string operations
case-sensitively.
Collations allow specifying alternative comparison rules. This is
useful for achieving, e.g., case- or accent-insensitive string
operations, or language-specific string sorting.
Collations are supported by many engines: Databricks
,
Spark

,
Snowflake ,
Oracle


 - to
name just a few - this list is not complete.

In Snowflake, we see heavy use of the collation feature. Several
users have approached us, mentioning they want to migrate to
Iceberg tables, but are currently blocked by Iceberg's lack of
collation support.

Given the widespread support for collations across different
engines, I believe introducing collations to Iceberg will increase
interoperability and boost its adoption.
I'd be curious about your thoughts.

*Goal of the proposal*
- Support collation specifications for columns
- Define how collation bounds should be stored - UTF-8 based
bounds are not useful for collated columns

*Required Changes*
- Extend the schema to let (string) fields be annotated with a
collation

More details can be found in this doc



Re: [Discussion] Collation Support

2026-03-28 Thread Andrei Tserakhau via dev
Hi Alexander,

This looks really interesting. We've been working on collation support in
Delta and have shipped it in production for some time, so this is an area
we care about a lot. If this proposal moves forward we'd be happy to
collaborate on the design and implementation.

The pseudo-field approach for collation metrics is clean and composes well
with existing Iceberg infrastructure. The specifier coverage is
comprehensive.

A few areas worth discussing as this evolves:

1 - Sort key stability and versioning

ICU sort keys are not stable across versions, so a pinned ICU version bump
in a future Iceberg release would invalidate all existing collation
metrics. In multi-engine environments, requiring all engines to converge on
one ICU version is unrealistic.

We store original string values instead of sort keys and allow per-file
version annotations -- worth discussing whether something similar could
work here.

2 - Provider abstraction

The proposal assumes ICU as the sole provider, but Spark ships non-ICU
collations like UTF8_LCASE that are widely used. A provider or namespace
layer would prevent name collisions and support engine-specific collations
without future spec changes.

3 - Operational surface

A few things that turned out correctness-critical in our implementation:
partition transforms on collated columns (collation-equal but byte-distinct
values in different directories), sort order semantics, equality deletes
under collation, and Parquet filter pushdown (must be disabled since
Parquet has no collation concept).

These don't all need to be solved in v1 but would help to scope them.

4 - Smaller items (nit's)

UTF-8 bounds for the original field id should be "must write" not "should"
-- otherwise backward compat breaks for non-aware engines. Engine fallback
behavior (case-sensitive vs older ICU vs fail) could use a recommended
preference order to avoid divergent results across engines. The collation
specifier syntax would benefit from a formal grammar.

---

Happy to share our Delta design doc and implementation learnings in more
detail. Looking forward to the discussion.

Best,
Andrei

On Sat, Mar 28, 2026 at 11:49 PM Alexander Löser 
wrote:

> Hi everyone,
>
> this is my first interaction with the Iceberg community, so here a few
> words about myself:
> - I'm Alex, a Berlin-based software engineer
> - I've been working at Snowflake for 4 years now
> - I spend most of my time on data types, particularly binary, strings and
> collations.
>
> I'd like to start a discussion about adding collations to the Iceberg spec.
>
> Conceptually, collations are an annotation on the string data type. By
> default, most engines perform string operations case-sensitively.
> Collations allow specifying alternative comparison rules. This is useful
> for achieving, e.g., case- or accent-insensitive string operations, or
> language-specific string sorting.
> Collations are supported by many engines: Databricks
> ,
> Spark
> ,
> Snowflake , Oracle
> 
>  - to
> name just a few - this list is not complete.
>
> In Snowflake, we see heavy use of the collation feature. Several users
> have approached us, mentioning they want to migrate to Iceberg tables, but
> are currently blocked by Iceberg's lack of collation support.
>
> Given the widespread support for collations across different engines, I
> believe introducing collations to Iceberg will increase interoperability
> and boost its adoption.
> I'd be curious about your thoughts.
>
> *Goal of the proposal*
> - Support collation specifications for columns
> - Define how collation bounds should be stored - UTF-8 based bounds are
> not useful for collated columns
>
> *Required Changes*
> - Extend the schema to let (string) fields be annotated with a collation
>
> More details can be found in this doc
> 
> .
>
> I'm also hoping to present the idea in the next community sync.
>
> Best, Alex
>
>
>


[Discussion] Collation Support

2026-03-28 Thread Alexander Löser

Hi everyone,

this is my first interaction with the Iceberg community, so here a few 
words about myself:

- I'm Alex, a Berlin-based software engineer
- I've been working at Snowflake for 4 years now
- I spend most of my time on data types, particularly binary, strings 
and collations.


I'd like to start a discussion about adding collations to the Iceberg spec.

Conceptually, collations are an annotation on the string data type. By 
default, most engines perform string operations case-sensitively.
Collations allow specifying alternative comparison rules. This is useful 
for achieving, e.g., case- or accent-insensitive string operations, or 
language-specific string sorting.
Collations are supported by many engines: Databricks 
, 
Spark 
, 
Snowflake , 
Oracle 
 - to 
name just a few - this list is not complete.


In Snowflake, we see heavy use of the collation feature. Several users 
have approached us, mentioning they want to migrate to Iceberg tables, 
but are currently blocked by Iceberg's lack of collation support.


Given the widespread support for collations across different engines, I 
believe introducing collations to Iceberg will increase interoperability 
and boost its adoption.

I'd be curious about your thoughts.

*Goal of the proposal*
- Support collation specifications for columns
- Define how collation bounds should be stored - UTF-8 based bounds are 
not useful for collated columns


*Required Changes*
- Extend the schema to let (string) fields be annotated with a collation

More details can be found in this doc 
.


I'm also hoping to present the idea in the next community sync.

Best, Alex