voonhous opened a new pull request, #20041: URL: https://github.com/apache/hudi/pull/20041
### Describe the issue this Pull Request addresses On Spark 4.1+ (`spark.sql.variant.pushVariantIntoScan` on by default) a query extracting several paths from a variant column throws `NegativeArraySizeException` once a row lacks enough of them, on any MOR read that merges a log and on any bootstrap read. Follow-up to #19783: two row writers are still typed from the engine schema, so the pushed projection struct is copied as a `VariantType` and `UnsafeRow.getVariant` reads its null bitset as the variant length; small projections survive by coincidence. Also answers the #19783 review question: the rule matches every Hudi read relation alike. CDC is untouched by construction (its relation schema is four strings) and `hoodie.datasource.read.use.new.parquet.file.format` no longer exists. ### Summary and Changelog - `BaseSparkInternalRecordContext.setRowShape` / `getRowStructType`: the Spark type rows carry in a read; `projectRecord` (the output converter) and `getBootstrapProjection` (the skeleton/data join) build their row writers over it. - `SparkFileFormatInternalRowReaderContext.setSchemaHandler` installs the PushVariantIntoScan overlay once the merger is known: base rows always carry it, log rows only when `shouldProjectVariants` rewrites them. - `TestVariantShreddingMixedLayouts`: time travel, incremental V1 and V2, METADATA_ONLY bootstrap (COW, MOR with a log) on both conf arms, each asserting whether the scan carries the projection struct, plus the eight-path query that failed before the fix. <details><summary>Details</summary> - Incremental V1 is reached through `hoodie.datasource.read.incr.table.version=6` and bounded at the base instant's requested time; the zero-row V2 read of the same bound proves the option took effect rather than falling through to V2. - The bootstrap legs run on the default record type: a METADATA_ONLY bootstrap on the SPARK record type fails in the skeleton-file write (HUDI-5807, `TestDataSourceForBootstrap` skips it too), independently of variants. The writer's column-stats index is disabled for the same reason that suite disables it (rejected by the bootstrap commit). - Payload-based tables with log files keep the previous behaviour: their log rows are not rewritten (#18674), so no single row shape is right for the output converter there. - Verified locally on `-Dspark4.1` (Spark 4.1.1): `TestVariantShreddingMixedLayouts`, `TestVariantDataType`, `TestBaseSpark4AdapterVariantMethods`, `TestStreamingSource`. CI runs the suites on spark4.2, where the conf default is the same. </details> ### Impact Reads that hit the exception now return the extracted values; no API or config change. ### Risk Level low: the overlay only changes the Spark type of a row writer when the query carries a projection struct, i.e. rows that already arrive in that shape. ### Documentation Update none ### Contributor's checklist - [x] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [x] Enough context is provided in the sections above - [x] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
