voonhous opened a new pull request, #20127:
URL: https://github.com/apache/hudi/pull/20127

   ### Describe the issue this Pull Request addresses
   
   Validation of Spark 4.1+ `PushVariantIntoScan` on master, schema-evolution 
item 3 of the checklist on #18285 
(https://github.com/apache/hudi/issues/18285#issuecomment-5869224297): an 
explicit `ALTER TABLE ... ALTER COLUMN n TYPE bigint` on a column beside a 
variant, without schema-on-read.
   
   Finding: the explicit path cannot widen anything. Without 
`hoodie.schema.on.read.enable`, `HoodieCatalog.loadTable` hands the table back 
as a V1Table and the Hudi ALTER rewrite rule is a no-op, so Spark's 
`ResolveSessionCatalog` turns `ALTER COLUMN ... TYPE` (and the Hive-style 
`CHANGE COLUMN`) into its v1 `AlterTableChangeColumnCommand`, which 
`HoodiePostAnalysisRule` swaps for `AlterHoodieTableChangeColumnCommand`. That 
command refuses every data-type change before `commitWithSchema`, and Spark's 
own up-cast check on `AlterColumns` never runs on this route, so re-typing the 
variant itself, or a sibling into a variant, meets the same refusal. A struct 
member never reaches Hudi: Spark refuses a qualified column on a v1 table at 
analysis. The DataFrame write pinned by #20125 and #20126 is therefore the only 
way a sibling gets widened beside a variant, and the "then the implicit read 
path" half of the checklist item has nothing new to read.
   
   ### Summary and Changelog
   
   Test only. `TestVariantShreddingMixedLayouts` gets "Explicit ALTER COLUMN 
TYPE beside a variant is refused without schema-on-read":
   
   - a top-level table `(id int, v variant, n int, ts long)` and a nested one 
`(id int, s struct<inner: variant, n: int>, ts long)`, COW and MOR, each seeded 
with a shredded base file and a second commit (a log block on MOR);
   - refused statements, one pinned message each: `ALTER COLUMN n TYPE bigint`, 
`CHANGE COLUMN n n bigint`, `ALTER COLUMN v TYPE string` and `ALTER COLUMN n 
TYPE variant` (Hudi: "ALTER TABLE CHANGE COLUMN is not supported for changing 
column ..."); `ALTER COLUMN s.n TYPE bigint` and `ALTER COLUMN s.inner TYPE 
string` (Spark: "does not support ALTER COLUMN with qualified column");
   - pins that nothing happened: no new instant, catalog types unchanged;
   - the variant projection still reads back afterwards on both 
`spark.sql.variant.pushVariantIntoScan` arms with the plan pinned per arm, and 
one more insert lands.
   
   Result: green on master as-is. No production change.
   
   ### Impact
   
   None.
   
   ### Risk Level
   
   none
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [x] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [x] Enough context is provided in the sections above
   - [x] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to