voonhous opened a new issue, #20141:
URL: https://github.com/apache/hudi/issues/20141

   ## Summary
   
   Renaming a VARIANT column does not work today on either schema path. This 
issue lists every failure on the way, with its error and what has to change, so 
that the pins landing under #18285 (checklist item 5, 
https://github.com/apache/hudi/issues/18285#issuecomment-5869224297) reference 
one place instead of explaining the mechanism in test comments.
   
   Everything below was run on Spark 4.1.1, Scala 2.13, JDK 17, master 
`9e9f7336a49e`, COW and MOR, SPARK record type, shredded and unshredded base 
files, table `(id int, v variant, ts long)` renamed `v` to `w`.
   
   ## Schema-on-read (the only path with a rename primitive)
   
   **1. The DDL is refused under default configs.**
   
   ```
   org.apache.hudi.exception.MissingSchemaFieldException: Schema validation 
failed due to missing field. Fields missing from incoming schema: {v}
   ```
   
   `AlterTableCommand.commitWithSchema` runs `HoodieTable.validateSchema` since 
#13595, and the writer-schema check sees the rename as a dropped field. Tracked 
by #19766. 
`hoodie.datasource.write.schema.allow.auto.evolution.column.drop=true` 
short-circuits the check and the DDL commits. A failed attempt leaves a 
REQUESTED instant on the timeline because `client.startCommit` runs before the 
validation.
   
   After the DDL commits: it is its own instant, no data file is rewritten, the 
footer still names `v`, and `describe` shows `w variant`.
   
   **2. Reads of the renamed column with 
`spark.sql.variant.pushVariantIntoScan` on (the default) fail through the 
schema-on-read guard.**
   
   ```
   org.apache.hudi.exception.HoodieException: Column 'w' is a variant requested 
in Spark's full-variant projection shape - by the PushVariantIntoScan rewrite 
(spark.sql.variant.pushVariantIntoScan) on a query, or by Hudi's own base-file 
reads for compaction, clustering and CDC - and the table is read with 
schema-on-read (hoodie.schema.on.read.enable), which cannot reconstruct 
variants (see issue #18285). ...
   ```
   
   `ParquetSchemaEvolutionUtils.validateNoShreddedVariants`, projection-shape 
arm. On COW it fires from the base-file reader, on MOR from 
`SparkFileFormatInternalRowReaderContext` through the file group reader. 
Unblocked by the #18285 design: the variant column has to pass through the 
internal-schema pipeline as an atomic token (field-id mapping at the column 
level, the catalyst type spliced verbatim into the merged request, no 
type-change entry, no filter remap inside the synthetic struct), after which 
the guard's rewrite arm is retired.
   
   **3. Reads with pushdown off die in generated code.**
   
   ```
   java.lang.ClassCastException: class 
org.apache.spark.sql.catalyst.expressions.SpecificInternalRow cannot be cast to 
class org.apache.spark.unsafe.types.VariantVal
       at 
org.apache.spark.sql.catalyst.expressions.BaseGenericInternalRow.getVariant(rows.scala:52)
       at 
org.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificUnsafeProjection.apply(Unknown
 Source)
   ```
   
   Same on every leg, shredded or not. The guard's shredded-file arm resolves 
footer columns by the query-schema name and skips a renamed column, as its 
scaladoc says. The merged internal-schema request then carries `w` as 
`struct<metadata, value>` because the internal schema has no VARIANT arm 
(`InternalSchemaConverter` round-trips the sentinel record, 
`SparkInternalSchemaConverter.constructSparkSchemaFromInternalSchema` does 
not), the reader hands back a struct row where the plan's output type is 
variant, and the generated projection's `getVariant` cast fails. Two things 
unblock it: a VARIANT arm in the Spark-side converter so the merged request 
keeps `VariantType`, and reconstruction of shredded files (`value` plus 
`typed_value`) under schema-on-read, which is #18285 itself. Until then the 
guard could resolve renamed columns by field id so this arm fails through the 
guard rather than in codegen; today it is loud by accident.
   
   **4. `count(*)` fails in the vectorized reader once any internal schema is 
committed.** Not rename-specific: #20139, with the session-conf side effect 
that hides it in #20140.
   
   **5. Reads without schema-on-read after the DDL.** `w` reads null on every 
row, which is the schema-on-read contract (reads without it resolve by name), 
but the V1 catalog schema now types `w` as `struct<metadata: binary, value: 
binary>`: the DDL rewrote the catalog through the same converter without a 
VARIANT arm. Unblocked by the converter arm in 3.
   
   ## Schema-on-write (no rename primitive)
   
   **1. `ALTER TABLE ... RENAME COLUMN` without schema-on-read is refused by 
Spark at analysis**, with or without the drop knob:
   
   ```
   org.apache.spark.sql.AnalysisException: 
[UNSUPPORTED_FEATURE.TABLE_OPERATION] The feature is not supported: Table 
`spark_catalog`.`default`.`t` does not support RENAME COLUMN.
   ```
   
   `HoodieCatalog.loadTable` hands back a V1 table when schema evolution is 
off, and V1 tables have no rename. Nothing to unblock short of a rename 
primitive, and field ids only exist under schema-on-read.
   
   **2. A DataFrame write carrying `w` instead of `v`** is refused under 
defaults:
   
   ```
   org.apache.hudi.exception.MissingSchemaFieldException: Schema validation 
failed due to missing field. Fields missing from incoming schema: {t_record.v}
   ```
   
   With `allow.auto.evolution.column.drop=true` it commits as drop `v` plus add 
`w`: the table schema becomes `(id, ts, w)`, rows written before it read null 
for both names, and the SQL catalog still says `v` (a path write does not 
update it), so `w` cannot be resolved through SQL at all. A pushdown-on read of 
`v` over the new file then hits #20135 (fixed by #20136). That is the knob's 
documented drop semantics, not a rename.
   
   Conclusion: rename support is a schema-on-read feature. On schema-on-write 
the only work is to make the drop-plus-add outcome and the catalog divergence 
explicit, or refuse it.
   
   ## What supported looks like
   
   - The DDL commits under default configs (#19766).
   - Reads under schema-on-read return the data on both pushdown arms, COW and 
MOR, shredded or not (#18285 design, converter VARIANT arm, guard retired).
   - The catalog keeps `variant` after any schema-on-read DDL.
   - Reads without schema-on-read are documented as name-resolved, like any 
other column.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to