ashokraminedi opened a new pull request, #20130:
URL: https://github.com/apache/hudi/pull/20130
### Describe the issue this Pull Request addresses
Closes #20108
Schema-on-read vectorized batch reads can return an incorrect value when an
INT/LONG column is evolved to a DECIMAL type and the existing value cannot fit
within the target decimal precision and scale.
For example, when an INT value `12345` is read after the column is evolved
to `DECIMAL(4, 2)`, the vectorized read path returns `123.45` instead of `NULL`
when ANSI mode is disabled.
The issue is in the INT/LONG to DECIMAL conversion path in
`SparkInternalSchemaConverter`. `Decimal.changePrecision()` returns a boolean
indicating whether the value can be represented using the requested precision
and scale. The existing implementation ignored this return value and wrote the
decimal to the output column vector even when `changePrecision()` returned
`false`.
This PR fixes the reproduced INT/LONG to DECIMAL precision-overflow case by
honoring the result of `changePrecision()` and writing `NULL` when the
conversion cannot be represented using the target precision and scale.
**Additional finding:**
During the investigation, I also noticed that some other source-type
conversions are not currently handled by the vectorized conversion logic,
including Boolean, Byte, Short, Binary, and Timestamp. These cases are outside
the scope of this narrow fix and should be validated separately to determine
whether additional conversions or an appropriate fallback path are needed.
### Summary and Changelog
- Check the return value of `Decimal.changePrecision()` during INT/LONG to
DECIMAL vectorized conversion.
- Write `NULL` to the output column vector when the source value cannot fit
the target DECIMAL precision and scale.
- Preserve the existing behavior when the decimal conversion succeeds.
- Add a regression test reproducing INT to `DECIMAL(4, 2)` precision
overflow with:
- schema-on-read enabled
- Parquet vectorized reader enabled
- ANSI mode disabled
- Run the regression for both COW and MOR table types.
### Impact
This fixes incorrect values returned by schema-on-read vectorized batch
reads for INT/LONG to DECIMAL conversions when the source value exceeds the
precision of the target DECIMAL type.
For the reproduced case:
`INT 12345 -> DECIMAL(4, 2)`
Before this change, the vectorized reader returned:
`123.45`
With this change, the reader returns:
`NULL`
when ANSI mode is disabled, instead of exposing the incorrectly rescaled
decimal value after `changePrecision()` fails.
The production change is limited to the existing INT/LONG to DECIMAL
conversion path in `SparkInternalSchemaConverter`.
There are no public API, configuration, or storage-format changes.
### Risk Level
low
The change is localized to handling the failure result from
`Decimal.changePrecision()` in the existing INT/LONG to DECIMAL vectorized
conversion path.
The regression was verified for both COW and MOR tables.
Before the production fix, the new tests reproduced the issue:
- 4 tests run
- 2 succeeded
- 2 failed
- Both new COW and MOR cases returned `123.45` instead of the expected `NULL`
After the fix:
- 4 tests run
- 4 succeeded
- 0 failed
Java Checkstyle for `hudi-spark-client` also completes with 0 violations.
### Documentation Update
none
This is a correctness fix to an existing schema-evolution conversion path
and does not introduce any new API, configuration, or user-facing feature.
### Contributor's checklist
- [X] Read through [contributor's
guide](https://hudi.apache.org/contribute/how-to-contribute)
- [X] Enough context is provided in the sections above
- [X] Adequate tests were added if applicable
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]