JunRuiLee opened a new pull request, #805:
URL: https://github.com/apache/paimon-rust/pull/805

   ### Purpose
   
   Linked issue: _TBD — filing one before this leaves draft._
   
   `ALTER TABLE t ADD COLUMN parent.child` bumps the table schema without 
rewriting data files, so a file written before the change holds a struct with 
fewer children than the read type declares. Reconciling that column ends at 
Arrow's `cast`, which pairs struct fields by position and requires equal arity, 
so the read fails outright instead of returning NULL for the added nested field:
   
   ```
   Failed to cast column 'media' from Struct([codec]) to Struct([codec, 
color_transfer]):
   Invalid argument error: Incorrect number of arrays for StructArray fields, 
expected 2 got 1
   ```
   
   Top-level schema evolution is already handled (the field-id index mapping 
null-fills a missing column); only evolution *inside* a ROW is affected.
   
   Java handles this in `SchemaEvolutionUtil.createRowCastExecutor`, which 
builds a per-ROW index mapping by field id and yields NULL for a target child 
with no source counterpart, with `createArrayCastExecutor` / 
`createMapCastExecutor` covering `ARRAY<ROW>` and `MAP<_, ROW>`. Worth noting 
for anyone who greps for this: the `// todo: support nested field missing` in 
`ParquetReaderFactory.clipParquetType` is on a path Java's own read never 
reaches — `FormatReaderMapping.pruneDataType` narrows the read type to what the 
file actually has first, and the row cast executor then supplies the NULLs.
   
   ### Brief change log
   
   - **`crates/paimon/src/arrow/nested_evolution.rs`** (new): `evolve_column` 
reconciles a decoded column with the read type — pair ROW children by **field 
id**, recurse so an added field is handled at any depth and inside ARRAY 
elements / MAP values, cast promoted leaves, and fill a child the file does not 
carry with NULLs. Pairing by id rather than by name matters twice: a renamed 
nested column still resolves, and a child dropped and re-added under the same 
name and type reads as NULL instead of serving the dropped field's values (the 
Arrow types are identical there, so an Arrow-level comparison cannot tell them 
apart). Variant-extraction rows are deliberately left to the cast path — their 
field ids are positional, not schema ids.
   - **`crates/paimon/src/table/data_file_reader.rs`**: both 
column-reconciliation sites (the main stream and `project_file_batch`) now 
share one `reconcile_column` helper. Each resolves its source field against the 
list actually handed to the format reader — the pruned fields for Parquet, the 
file's own fields for `.row` — so the source type describes what came back. 
This also fixes the no-evolution fallback to resolve by field id against the 
table schema, since `read_type` may be a nested projection of it.
   - **`crates/paimon/src/spec/types.rs`**: `DataType::equals_ignore_nullable`, 
mirroring Java `DataType.equalsIgnoreNullable`, used as the fast-path guard in 
`evolve_column` the way Java uses it in `createCastExecutor`.
   
   A ROW column whose decoded struct is missing a child the file's own schema 
declares now fails with `DataInvalid` rather than reading as NULL. That state 
means the file, its schema, or the reader's projection disagree; NULL-filling 
it would pass the gap off as legitimately absent data. Java cannot reach this 
case at all, because it addresses nested fields positionally rather than by 
name.
   
   ### Tests
   
   - `crates/paimon/tests/nested_schema_evolution_test.rs` (new, end-to-end): 
builds a primary-key table, commits one snapshot under a schema whose `media` 
ROW has a single child, then adds `media.color_transfer` as a new schema 
version and commits a second snapshot. A full read must return the older row 
with `color_transfer` NULL and its sibling value intact. Against `main` at 
`78994db` this test fails with the cast error quoted above.
   - 11 unit tests in `nested_evolution.rs`: NULL-fill for an added child; 
pairing by id rather than name (renamed nested column); dropping children the 
read type does not ask for; recursion into a nested ROW; preserving the 
struct's row-level validity buffer; casting a promoted nested leaf; added child 
inside an ARRAY element and inside a MAP value; the dropped-and-re-added case; 
the missing-child failure; and the untouched-when-types-match fast path.
   - `cargo test -p paimon`: 2631 lib tests + all test binaries pass.
   - `cargo fmt --all -- --check` and `cargo clippy --locked --all-targets 
--workspace --features fulltext,vortex -- -D warnings` are clean (locally 
excluding `pypaimon_rust`, whose `abi3-py310` needs a newer Python than this 
machine has).
   
   ### API and Format
   
   No storage format change, and no change to how a schema is written — this is 
read-side only. One new public method, `DataType::equals_ignore_nullable`; 
everything else added is `pub(crate)`.
   
   ### Documentation
   
   None needed.
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to