voonhous opened a new issue, #20139:
URL: https://github.com/apache/hudi/issues/20139

   ## Bug Description
   
   **What happened:**
   
   On a table with a VARIANT column read with 
`hoodie.schema.on.read.enable=true`, `select count(*)` fails once any 
schema-on-read DDL has committed an internal schema (an `add columns` is 
enough). Spark 4.1.1, master `9e9f7336a49e`, COW, SPARK record type, shredded 
or unshredded file:
   
   ```
   org.apache.spark.SparkException: [FAILED_READ_FILE.NO_HINT] Encountered 
error while reading file ...
   Caused by: java.lang.NullPointerException
       at 
org.apache.spark.sql.execution.datasources.parquet.HoodieVectorizedParquetRecordReader.close(HoodieVectorizedParquetRecordReader.java:97)
       at 
org.apache.spark.sql.execution.datasources.RecordReaderIterator.close(RecordReaderIterator.scala:69)
       at 
org.apache.spark.sql.execution.datasources.parquet.Spark41ParquetReader.buildVectorizedIterator(Spark41ParquetReader.scala:251)
   ```
   
   The NPE masks the real failure. With a null guard on `close`, the exception 
is Spark's:
   
   ```
   org.apache.spark.sql.AnalysisException: Invalid Spark read type: expected 
optional group v (VARIANT(1)) { required binary metadata; optional binary 
value; } to be variant type but found STRUCT<metadata: BINARY NOT NULL, value: 
BINARY NOT NULL>
       at 
org.apache.spark.sql.execution.datasources.parquet.ParquetSchemaConverter$.checkConversionRequirement(ParquetSchemaConverter.scala:897)
       at 
org.apache.spark.sql.execution.datasources.parquet.ParquetToSparkSchemaConverter.convertGroupField(ParquetSchemaConverter.scala:383)
       ...
       at 
org.apache.spark.sql.execution.datasources.parquet.VectorizedParquetRecordReader.initialize(VectorizedParquetRecordReader.java:199)
   ```
   
   **To reproduce** (Spark SQL, Spark 4.1):
   
   ```sql
   create table t (id int, v variant, ts long) using hudi
     tblproperties (primaryKey = 'id', preCombineField = 'ts', type = 'cow') 
location '/tmp/t';
   insert into t values (1, parse_json('{"a":1}'), 1000);
   set hoodie.schema.on.read.enable=true;
   alter table t add columns (note string);
   select count(*) from t;   -- fails as above in a fresh session
   ```
   
   **Cause:**
   
   Two parts.
   
   1. `ParquetSchemaEvolutionUtils.getHadoopConfClone` skips the 
shredded-variant guard for an empty projection (`requiredSchema.isEmpty`) but 
still merges the file schema with the UNPRUNED query schema and sets the result 
as `SPARK_ROW_REQUESTED_SCHEMA`. The internal schema has no VARIANT arm, so the 
variant column arrives as `struct<metadata, value>`. Spark 4.1's vectorized 
reader (`ParquetToSparkSchemaConverter.checkConversionRequirement`) refuses a 
VARIANT-annotated parquet group whose requested Spark type is not variant. The 
row-based reader accepts the same request 
(`spark.sql.parquet.enableVectorizedReader=false` makes the count pass).
   2. `HoodieVectorizedParquetRecordReader.close` iterates `idToColumnVectors`, 
which `initBatch` creates. `Spark41ParquetReader.buildVectorizedIterator` 
closes the iterator when `initialize` throws, before `initBatch` ran, so the 
close NPEs and the original exception is lost.
   
   **Why the existing test does not catch it:**
   
   `TestVariantShreddingMixedLayouts."Schema-on-read reads of shredded variant 
files fail fast"` pins `count(*)` after the DDL and is green, but only because 
the two variant reads before it made 
`HoodieFileGroupReaderBasedFileFormat.buildReaderWithPartitionValues` write 
`spark.sql.parquet.enableVectorizedReader=false` into the session conf (see the 
companion issue on that session-conf write). Run the count as the first read on 
a fresh session and it fails.
   
   **Expected behavior:**
   
   An empty projection reads no column data, so the requested parquet schema 
for it should be empty (or pruned to the columns actually read), not the whole 
merged table schema; and `close` should tolerate a reader whose `initBatch` 
never ran so the real exception surfaces.
   
   ## Environment
   
   Spark 4.1.1, Scala 2.13, JDK 17, local filesystem, master `9e9f7336a49e`. 
Found while probing rename under schema-on-read for #18285 (checklist item 5).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to