ashokraminedi opened a new pull request, #20130:
URL: https://github.com/apache/hudi/pull/20130

   ### Describe the issue this Pull Request addresses
   
   Closes #20108
   
   Schema-on-read vectorized batch reads can return an incorrect value when an 
INT/LONG column is evolved to a DECIMAL type and the existing value cannot fit 
within the target decimal precision and scale.
   
   For example, when an INT value `12345` is read after the column is evolved 
to `DECIMAL(4, 2)`, the vectorized read path returns `123.45` instead of `NULL` 
when ANSI mode is disabled.
   
   The issue is in the INT/LONG to DECIMAL conversion path in 
`SparkInternalSchemaConverter`. `Decimal.changePrecision()` returns a boolean 
indicating whether the value can be represented using the requested precision 
and scale. The existing implementation ignored this return value and wrote the 
decimal to the output column vector even when `changePrecision()` returned 
`false`.
   
   This PR fixes the reproduced INT/LONG to DECIMAL precision-overflow case by 
honoring the result of `changePrecision()` and writing `NULL` when the 
conversion cannot be represented using the target precision and scale.
   
   **Additional finding:** 
   During the investigation, I also noticed that some other source-type 
conversions are not currently handled by the vectorized conversion logic, 
including Boolean, Byte, Short, Binary, and Timestamp. These cases are outside 
the scope of this narrow fix and should be validated separately to determine 
whether additional conversions or an appropriate fallback path are needed.
   
   ### Summary and Changelog
   
   - Check the return value of `Decimal.changePrecision()` during INT/LONG to 
DECIMAL vectorized conversion.
   - Write `NULL` to the output column vector when the source value cannot fit 
the target DECIMAL precision and scale.
   - Preserve the existing behavior when the decimal conversion succeeds.
   - Add a regression test reproducing INT to `DECIMAL(4, 2)` precision 
overflow with:
     - schema-on-read enabled
     - Parquet vectorized reader enabled
     - ANSI mode disabled
   - Run the regression for both COW and MOR table types.
   
   ### Impact
   
   This fixes incorrect values returned by schema-on-read vectorized batch 
reads for INT/LONG to DECIMAL conversions when the source value exceeds the 
precision of the target DECIMAL type.
   
   For the reproduced case:
   
   `INT 12345 -> DECIMAL(4, 2)`
   
   Before this change, the vectorized reader returned:
   
   `123.45`
   
   With this change, the reader returns:
   
   `NULL`
   
   when ANSI mode is disabled, instead of exposing the incorrectly rescaled 
decimal value after `changePrecision()` fails.
   
   The production change is limited to the existing INT/LONG to DECIMAL 
conversion path in `SparkInternalSchemaConverter`.
   
   There are no public API, configuration, or storage-format changes.
   
   ### Risk Level
   
   low
   
   The change is localized to handling the failure result from 
`Decimal.changePrecision()` in the existing INT/LONG to DECIMAL vectorized 
conversion path.
   
   The regression was verified for both COW and MOR tables.
   
   Before the production fix, the new tests reproduced the issue:
   
   - 4 tests run
   - 2 succeeded
   - 2 failed
   - Both new COW and MOR cases returned `123.45` instead of the expected `NULL`
   
   After the fix:
   
   - 4 tests run
   - 4 succeeded
   - 0 failed
   
   Java Checkstyle for `hudi-spark-client` also completes with 0 violations.
   
   ### Documentation Update
   
   none
   
   This is a correctness fix to an existing schema-evolution conversion path 
and does not introduce any new API, configuration, or user-facing feature.
   
   
   ### Contributor's checklist
   
   - [X] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [X] Enough context is provided in the sections above
   - [X] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to