jackylee-ch opened a new pull request, #10124:
URL: https://github.com/apache/paimon/pull/10124

   ### Purpose
   
   Reading and writing a `DECIMAL` column in the `row` file format both use 
Python decimal arithmetic under the **default** context (28 significant 
digits). For `DECIMAL(p)` with `p > 28`, the unscaled integer has up to 38 
digits, so:
   
   - the writer's `int(value * (10 ** scale))` rounds to 28 digits before 
encoding, and
   - the reader's `Decimal(unscaled) / Decimal(10 ** scale)` rounds to 28 
digits after decoding.
   
   A `DECIMAL(38, 0)` value `12345678901234567890123456789012345678` 
round-trips as `12345678901234567890123456790000000000` — silent data 
corruption.
   
   The fix computes the unscaled value (write) and rescales it (read) under a 
context wide enough for the column — `max(precision + abs(scale), 38)` with 
`scaleb` — mirroring the existing `pypaimon/data/decimal.py`. `DECIMAL(p <= 
18)` is unchanged.
   
   ### Tests
   
   `test_format_row_reader_writer.test_high_precision_decimal` round-trips 
positive, negative and null values across `DECIMAL(38,0)`, `DECIMAL(38,10)`, 
`DECIMAL(28,4)` and `DECIMAL(18,2)` and asserts exact values. Fails on master 
(`DECIMAL(38,0)` reads back rounded to 28 digits); passes here.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to