voonhous opened a new pull request, #20127: URL: https://github.com/apache/hudi/pull/20127
### Describe the issue this Pull Request addresses Validation of Spark 4.1+ `PushVariantIntoScan` on master, schema-evolution item 3 of the checklist on #18285 (https://github.com/apache/hudi/issues/18285#issuecomment-5869224297): an explicit `ALTER TABLE ... ALTER COLUMN n TYPE bigint` on a column beside a variant, without schema-on-read. Finding: the explicit path cannot widen anything. Without `hoodie.schema.on.read.enable`, `HoodieCatalog.loadTable` hands the table back as a V1Table and the Hudi ALTER rewrite rule is a no-op, so Spark's `ResolveSessionCatalog` turns `ALTER COLUMN ... TYPE` (and the Hive-style `CHANGE COLUMN`) into its v1 `AlterTableChangeColumnCommand`, which `HoodiePostAnalysisRule` swaps for `AlterHoodieTableChangeColumnCommand`. That command refuses every data-type change before `commitWithSchema`, and Spark's own up-cast check on `AlterColumns` never runs on this route, so re-typing the variant itself, or a sibling into a variant, meets the same refusal. A struct member never reaches Hudi: Spark refuses a qualified column on a v1 table at analysis. The DataFrame write pinned by #20125 and #20126 is therefore the only way a sibling gets widened beside a variant, and the "then the implicit read path" half of the checklist item has nothing new to read. ### Summary and Changelog Test only. `TestVariantShreddingMixedLayouts` gets "Explicit ALTER COLUMN TYPE beside a variant is refused without schema-on-read": - a top-level table `(id int, v variant, n int, ts long)` and a nested one `(id int, s struct<inner: variant, n: int>, ts long)`, COW and MOR, each seeded with a shredded base file and a second commit (a log block on MOR); - refused statements, one pinned message each: `ALTER COLUMN n TYPE bigint`, `CHANGE COLUMN n n bigint`, `ALTER COLUMN v TYPE string` and `ALTER COLUMN n TYPE variant` (Hudi: "ALTER TABLE CHANGE COLUMN is not supported for changing column ..."); `ALTER COLUMN s.n TYPE bigint` and `ALTER COLUMN s.inner TYPE string` (Spark: "does not support ALTER COLUMN with qualified column"); - pins that nothing happened: no new instant, catalog types unchanged; - the variant projection still reads back afterwards on both `spark.sql.variant.pushVariantIntoScan` arms with the plan pinned per arm, and one more insert lands. Result: green on master as-is. No production change. ### Impact None. ### Risk Level none ### Documentation Update none ### Contributor's checklist - [x] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [x] Enough context is provided in the sections above - [x] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
