Hi everyone,

I'd like to propose a clarification to the spec about how readers handle
data file columns whose type is narrower than the field's type in the table
schema.

*Background*

Type promotion lets a field evolve, for example from DECIMAL(8. 0) to
DECIMAL(25. 0) without rewriting data files. Existing Parquet files keep
storing the column as int32 with DECIMAL(8. 0) and readers must read them
as DECIMAL(25, 0).

The spec addresses reading narrower types in the context of schema
evolution. However, schema evolution can potentially leave no markers in
the metadata, so readers must be prepared to read narrower types than the
type specified in the schema. We have come across systems which do not read
all possible allowed types for columns but instead check for an exact type
match.

*Proposed Change*
PR: https://github.com/apache/iceberg/pull/18404

Add one sentence to Schema Evolution, immediately after the table of valid
type promotions:

*"Readers must accept a data file column whose type is either the field's
type or a type that can be promoted to the field's type according to the
valid type promotions for the table's format version, regardless of whether
the field's type was evolved."*

Please let me know what you think. Feedback is welcome here or on the PR.

Thank you
Sandeep Gottimukkala

Reply via email to