voonhous opened a new pull request, #20041:
URL: https://github.com/apache/hudi/pull/20041

   ### Describe the issue this Pull Request addresses
   
   On Spark 4.1+ (`spark.sql.variant.pushVariantIntoScan` on by default) a 
query extracting several paths from a variant column throws 
`NegativeArraySizeException` once a row lacks enough of them, on any MOR read 
that merges a log and on any bootstrap read. Follow-up to #19783: two row 
writers are still typed from the engine schema, so the pushed projection struct 
is copied as a `VariantType` and `UnsafeRow.getVariant` reads its null bitset 
as the variant length; small projections survive by coincidence.
   
   Also answers the #19783 review question: the rule matches every Hudi read 
relation alike. CDC is untouched by construction (its relation schema is four 
strings) and `hoodie.datasource.read.use.new.parquet.file.format` no longer 
exists.
   
   ### Summary and Changelog
   
   - `BaseSparkInternalRecordContext.setRowShape` / `getRowStructType`: the 
Spark type rows carry in a read; `projectRecord` (the output converter) and 
`getBootstrapProjection` (the skeleton/data join) build their row writers over 
it.
   - `SparkFileFormatInternalRowReaderContext.setSchemaHandler` installs the 
PushVariantIntoScan overlay once the merger is known: base rows always carry 
it, log rows only when `shouldProjectVariants` rewrites them.
   - `TestVariantShreddingMixedLayouts`: time travel, incremental V1 and V2, 
METADATA_ONLY bootstrap (COW, MOR with a log) on both conf arms, each asserting 
whether the scan carries the projection struct, plus the eight-path query that 
failed before the fix.
   
   <details><summary>Details</summary>
   
   - Incremental V1 is reached through 
`hoodie.datasource.read.incr.table.version=6` and bounded at the base instant's 
requested time; the zero-row V2 read of the same bound proves the option took 
effect rather than falling through to V2.
   - The bootstrap legs run on the default record type: a METADATA_ONLY 
bootstrap on the SPARK record type fails in the skeleton-file write (HUDI-5807, 
`TestDataSourceForBootstrap` skips it too), independently of variants. The 
writer's column-stats index is disabled for the same reason that suite disables 
it (rejected by the bootstrap commit).
   - Payload-based tables with log files keep the previous behaviour: their log 
rows are not rewritten (#18674), so no single row shape is right for the output 
converter there.
   - Verified locally on `-Dspark4.1` (Spark 4.1.1): 
`TestVariantShreddingMixedLayouts`, `TestVariantDataType`, 
`TestBaseSpark4AdapterVariantMethods`, `TestStreamingSource`. CI runs the 
suites on spark4.2, where the conf default is the same.
   </details>
   
   ### Impact
   
   Reads that hit the exception now return the extracted values; no API or 
config change.
   
   ### Risk Level
   
   low: the overlay only changes the Spark type of a row writer when the query 
carries a projection struct, i.e. rows that already arrive in that shape.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [x] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [x] Enough context is provided in the sections above
   - [x] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to