JunRuiLee opened a new pull request, #805: URL: https://github.com/apache/paimon-rust/pull/805
### Purpose Linked issue: _TBD — filing one before this leaves draft._ `ALTER TABLE t ADD COLUMN parent.child` bumps the table schema without rewriting data files, so a file written before the change holds a struct with fewer children than the read type declares. Reconciling that column ends at Arrow's `cast`, which pairs struct fields by position and requires equal arity, so the read fails outright instead of returning NULL for the added nested field: ``` Failed to cast column 'media' from Struct([codec]) to Struct([codec, color_transfer]): Invalid argument error: Incorrect number of arrays for StructArray fields, expected 2 got 1 ``` Top-level schema evolution is already handled (the field-id index mapping null-fills a missing column); only evolution *inside* a ROW is affected. Java handles this in `SchemaEvolutionUtil.createRowCastExecutor`, which builds a per-ROW index mapping by field id and yields NULL for a target child with no source counterpart, with `createArrayCastExecutor` / `createMapCastExecutor` covering `ARRAY<ROW>` and `MAP<_, ROW>`. Worth noting for anyone who greps for this: the `// todo: support nested field missing` in `ParquetReaderFactory.clipParquetType` is on a path Java's own read never reaches — `FormatReaderMapping.pruneDataType` narrows the read type to what the file actually has first, and the row cast executor then supplies the NULLs. ### Brief change log - **`crates/paimon/src/arrow/nested_evolution.rs`** (new): `evolve_column` reconciles a decoded column with the read type — pair ROW children by **field id**, recurse so an added field is handled at any depth and inside ARRAY elements / MAP values, cast promoted leaves, and fill a child the file does not carry with NULLs. Pairing by id rather than by name matters twice: a renamed nested column still resolves, and a child dropped and re-added under the same name and type reads as NULL instead of serving the dropped field's values (the Arrow types are identical there, so an Arrow-level comparison cannot tell them apart). Variant-extraction rows are deliberately left to the cast path — their field ids are positional, not schema ids. - **`crates/paimon/src/table/data_file_reader.rs`**: both column-reconciliation sites (the main stream and `project_file_batch`) now share one `reconcile_column` helper. Each resolves its source field against the list actually handed to the format reader — the pruned fields for Parquet, the file's own fields for `.row` — so the source type describes what came back. This also fixes the no-evolution fallback to resolve by field id against the table schema, since `read_type` may be a nested projection of it. - **`crates/paimon/src/spec/types.rs`**: `DataType::equals_ignore_nullable`, mirroring Java `DataType.equalsIgnoreNullable`, used as the fast-path guard in `evolve_column` the way Java uses it in `createCastExecutor`. A ROW column whose decoded struct is missing a child the file's own schema declares now fails with `DataInvalid` rather than reading as NULL. That state means the file, its schema, or the reader's projection disagree; NULL-filling it would pass the gap off as legitimately absent data. Java cannot reach this case at all, because it addresses nested fields positionally rather than by name. ### Tests - `crates/paimon/tests/nested_schema_evolution_test.rs` (new, end-to-end): builds a primary-key table, commits one snapshot under a schema whose `media` ROW has a single child, then adds `media.color_transfer` as a new schema version and commits a second snapshot. A full read must return the older row with `color_transfer` NULL and its sibling value intact. Against `main` at `78994db` this test fails with the cast error quoted above. - 11 unit tests in `nested_evolution.rs`: NULL-fill for an added child; pairing by id rather than name (renamed nested column); dropping children the read type does not ask for; recursion into a nested ROW; preserving the struct's row-level validity buffer; casting a promoted nested leaf; added child inside an ARRAY element and inside a MAP value; the dropped-and-re-added case; the missing-child failure; and the untouched-when-types-match fast path. - `cargo test -p paimon`: 2631 lib tests + all test binaries pass. - `cargo fmt --all -- --check` and `cargo clippy --locked --all-targets --workspace --features fulltext,vortex -- -D warnings` are clean (locally excluding `pypaimon_rust`, whose `abi3-py310` needs a newer Python than this machine has). ### API and Format No storage format change, and no change to how a schema is written — this is read-side only. One new public method, `DataType::equals_ignore_nullable`; everything else added is `pub(crate)`. ### Documentation None needed. 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
