wombatu-kun commented on code in PR #20136:
URL: https://github.com/apache/hudi/pull/20136#discussion_r4140463292
##########
hudi-client/hudi-spark-client/src/main/scala/org/apache/spark/sql/execution/datasources/parquet/HoodieParquetReadSupport.scala:
##########
@@ -49,7 +49,11 @@ class HoodieParquetReadSupport(
} else {
readContext.getRequestedSchema
}
- val trimmedParquetSchema =
HoodieParquetReadSupport.trimParquetSchema(requestedParquetSchema,
context.getFileSchema)
+ // Same condition as Spark's own intersectParquetGroups: only the
row-based reader wants a
+ // requested schema restricted to what the file has; the vectorized reader
null-fills a missing
+ // column itself and matches columns by position.
+ val trimmedParquetSchema =
HoodieParquetReadSupport.trimParquetSchema(requestedParquetSchema,
+ context.getFileSchema, dropMissingTopLevelFields =
!enableVectorizedReader)
Review Comment:
`HoodieSparkParquetReader.getUnsafeRowIterator` builds this read support
with `enableVectorizedReader = true` for its row-based `ParquetReader`, so a
parquet log block written before the `ADD COLUMNS` still keeps the missing `v2`
and a pushed-down read of it should hit the same
`INVALID_VARIANT_SHREDDING_SCHEMA`. No vectorized read ever goes through
`HoodieParquetReadSupport` (every reader pins `READ_SUPPORT_CLASS` to Spark's
`ParquetReadSupport`), so could the top-level drop be unconditional like the
nested one, with an update before the DDL in the parquet-log MOR leg to cover
it?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]