voonhous commented on issue #18285:
URL: https://github.com/apache/hudi/issues/18285#issuecomment-5869219005

   Scope update after checking what the other formats and Spark allow.
   
   Type promotion is out of scope for a VARIANT column. No format defines a 
promotion to or from variant: Iceberg v3's promotion table lists unknown, int, 
date, float and decimal only; Delta's `typeWidening` covers integer, float, 
date and decimal widening; Paimon's cast matrix gives VARIANT only the identity 
rule, which is what its `ALTER TABLE ... MODIFY` gates on. Spark refuses it 
before any catalog sees it: `UpCastRule.canUpCast` has `case (VariantType, _) 
=> false` and no arm accepts a variant target. Hudi refuses it on both of its 
paths as well: `HoodieSchemaCompatibility` treats VARIANT against any other 
type as a mismatch (the `ALTER COLUMN TYPE` path without schema-on-read), and 
`SchemaChangeUtils.isTypeUpdateAllow` throws on nested types, which is how the 
internal schema models a variant.
   
   Evolution inside the value is what the type is for and needs no DDL. 
Different shredding schemas across files are a reader reconstruction concern 
and already work without schema-on-read.
   
   So the operations this ticket covers, all under 
`hoodie.schema.on.read.enable`, are:
   
   - [ ] add a VARIANT column to an existing table (top-level, and nested in a 
struct)
   - [ ] drop a VARIANT column
   - [ ] rename a VARIANT column
   - [ ] reposition a VARIANT column
   
   Each on COW and on MOR with log files, 
`spark.sql.variant.pushVariantIntoScan` on and off, with a shredded file in the 
mix.
   
   Current state on master and what blocks the above:
   
   - `ParquetSchemaEvolutionUtils.validateNoShreddedVariants` fails fast on any 
schema-on-read read that carries the PushVariantIntoScan projection struct or 
meets a shredded file; its message points here. Rename is not fronted by it 
(the footer lookup is by query name) and dies in Spark codegen instead.
   - `SparkInternalSchemaConverter.constructSparkSchemaFromInternalSchema` has 
no VARIANT arm, so the first schema-on-read DDL rewrites the catalog column to 
`struct<metadata, value>` and plain reads fail afterwards. This is the first 
fix.
   - Under schema-on-read the merged request materializes the variant as 
`{metadata, value}` while the plan expects `VariantType` or the projection 
struct. The internal-schema pipeline has to pass the variant column through as 
one atomic unit (prune, merge, type-change info, filter rebuild), replacing the 
guard's rewrite arm.
   - `FileGroupRecordBuffer.composeEvolvedSchemaTransformer` ordering, see the 
section above.
   - #19766: rename and drop fail for every column type under default configs 
since #13595. The rename and drop halves cannot be exercised until that is 
fixed or the tests set 
`hoodie.datasource.write.schema.allow.auto.evolution.column.drop=true`.
   
   No test on master adds, drops, renames or repositions a VARIANT column 
through DDL. The existing evolution legs add a string sibling only.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to