voonhous opened a new issue, #20135:
URL: https://github.com/apache/hudi/issues/20135
### Describe the problem you faced
On Spark 4.1 with `spark.sql.variant.pushVariantIntoScan` on (the default),
adding a VARIANT column by DDL without schema-on-read breaks every Hudi
row-based parquet read of the files written before the DDL:
- any pushed-down read of a pre-DDL base file (`select variant_get(v2,
...)`, `cast(v2 as string)`, `v2 is null`) fails;
- a COW update of a pre-DDL file group through the SPARK-record merge handle
fails;
- MOR compaction of a pre-DDL file group fails, with parquet and with avro
log blocks.
Reads with pushdown off work, and the COW update through the AVRO-record
merge handle works.
### To Reproduce
```sql
create table t (id int, v variant, ts long) using hudi tblproperties
(primaryKey = 'id', preCombineField = 'ts') location '...';
insert into t select 1, parse_json('{"k":"a"}'), 1000;
alter table t add columns (v2 variant);
select id, variant_get(v2, '$.k', 'string') from t; -- fails
```
```
SparkException: [FAILED_READ_FILE.NO_HINT] ... <pre-DDL base>.parquet
Caused by: AnalysisException: [INVALID_VARIANT_SHREDDING_SCHEMA] The schema
`"STRUCT<`0`: STRING>"` is not a valid variant shredding schema.
at SparkShreddingUtils$.buildVariantSchema(SparkShreddingUtils.scala:561)
at
ParquetRowConverter$ParquetVariantConverter.<init>(ParquetRowConverter.scala:941)
at ParquetRowConverter.newConverter(ParquetRowConverter.scala:526)
at ParquetReadSupport.prepareForRead(ParquetReadSupport.scala:111)
at
Spark41ParquetReader.buildRowBasedIterator(Spark41ParquetReader.scala:290)
at SparkParquetReaderBase.read(SparkParquetReaderBase.scala:84)
at
HoodieFileGroupReaderBasedFileFormat.readBaseFile(HoodieFileGroupReaderBasedFileFormat.scala:625)
```
The write-side failures (COW merge handle with SPARK records, MOR
compaction) show the same error with `STRUCT<0: STRUCT<value: BINARY NOT NULL,
metadata: BINARY NOT NULL>>`, reached through
`SparkFileFormatInternalRowReaderContext.getFileRecordIterator` and
`FileGroupReaderBasedMergeHandle.doMerge` / `HoodieCompactor`.
### Root cause
Hudi's readers force `spark.sql.optimizer.nestedSchemaPruning.enabled=false`
(`SparkParquetReaderBase.read`,
`SparkReaderContextFactory.getHadoopConfiguration`). With pruning off, Spark's
`ParquetReadSupport.getRequestedSchema` skips `intersectParquetGroups`, so a
requested column the file does not have stays in the parquet requested schema
as a group synthesised from its catalyst type.
`HoodieParquetReadSupport.trimParquetSchema` then drops missing NESTED fields
but keeps a missing TOP-LEVEL field as it is. For a variant column requested as
the PushVariantIntoScan projection struct (or as the full-variant struct of the
internal reader), that synthesised group is a plain struct,
`ParquetRowConverter` still routes it to `ParquetVariantConverter`, and
`buildVariantSchema` rejects it. A plain `VariantType` converts to a real
variant group, which is why pushdown-off reads survive.
Vanilla Spark 4.1 reproduces the same error when reading a file lacking a
variant column with the vectorized reader off and nested schema pruning off,
and returns nulls with pruning on.
### Expected behavior
The pre-DDL files read `v2` as null on both pushdown arms, and the update
and compaction paths add a null `v2` to the rewritten rows.
### Environment Description
* Hudi version : master (9e9f7336a49e)
* Spark version : 4.1.1
* Storage : local
### Additional context
Found while validating `PushVariantIntoScan` schema evolution on master
(#18285 checklist, item "add a second VARIANT column by DDL"). Candidate fix:
drop a missing top-level field from the parquet read schema in
`trimParquetSchema` the way a missing nested one already is; Spark's row
converter builds converters per parquet field and leaves the catalyst columns
it never sees null, which is what Spark itself does through
`intersectParquetGroups` when pruning is on.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]