Hi everyone, I'd like to propose a clarification to the spec about how readers handle data file columns whose type is narrower than the field's type in the table schema.
*Background* Type promotion lets a field evolve, for example from DECIMAL(8. 0) to DECIMAL(25. 0) without rewriting data files. Existing Parquet files keep storing the column as int32 with DECIMAL(8. 0) and readers must read them as DECIMAL(25, 0). The spec addresses reading narrower types in the context of schema evolution. However, schema evolution can potentially leave no markers in the metadata, so readers must be prepared to read narrower types than the type specified in the schema. We have come across systems which do not read all possible allowed types for columns but instead check for an exact type match. *Proposed Change* PR: https://github.com/apache/iceberg/pull/18404 Add one sentence to Schema Evolution, immediately after the table of valid type promotions: *"Readers must accept a data file column whose type is either the field's type or a type that can be promoted to the field's type according to the valid type promotions for the table's format version, regardless of whether the field's type was evolved."* Please let me know what you think. Feedback is welcome here or on the PR. Thank you Sandeep Gottimukkala
