osipovartem opened a new pull request, #11177:
URL: https://github.com/apache/arrow-rs/pull/11177

   ## Which issue does this PR close?
   
   - Closes #11169.
   
   ## What changes are included in this PR?
   
   `JsonToVariant::append_json` now parses JSON directly into 
`VariantBuilderExt` while retaining exact numeric lexemes. The parser:
   
   - preserves the existing smallest-width integer encoding for values fitting 
in `i64`;
   - encodes fixed-point values and wider integers as the smallest fitting 
`Decimal4`, `Decimal8`, or `Decimal16`;
   - keeps exponent notation and values outside the Variant decimal range as 
`Double`;
   - applies the same rules recursively in arrays and objects;
   - borrows unescaped strings directly from the input and only allocates when 
JSON escapes must be decoded;
   - bounds nesting depth and does not commit a root value until the full input 
has been validated.
   
   This deliberately avoids enabling serde_json's workspace-wide 
`arbitrary_precision` feature and removes the intermediate `serde_json::Value` 
tree from the string conversion path.
   
   The first commit adds a defaulted fallible 
`VariantBuilderExt::try_append_value` method. This lets strict object builders 
report duplicate fields as `ArrowError` instead of panicking. It also fixes 
duplicate validation to check before mutating an object's field mapping, so the 
same builder remains usable after an error.
   
   The existing `append_json(&serde_json::Value, ...)` API remains available. 
Its documentation now explains that non-integer numbers use `Double` because 
the original fixed-point lexeme is no longer available.
   
   ## Are these changes tested?
   
   Yes. This enables the existing Decimal4/8/16 tests and adds coverage for 
nested decimals, precision/scale boundaries, malformed JSON, Unicode escapes 
and surrogate pairs, duplicate fields, recursion limits, error rollback, and 
the documented difference between the string and `Value` APIs.
   
   Local validation:
   
   - `cargo +1.95.0 test -p parquet-variant --lib` (175 passed)
   - `cargo +1.95.0 test -p parquet-variant-json` (76 unit tests and 6 doctests 
passed)
   - `cargo +1.95.0 test -p parquet-variant-compute --lib` (367 passed)
   - `cargo +1.95.0 clippy -p parquet-variant-json --all-targets -- -D warnings`
   - `cargo +1.89.0 check -p parquet-variant-json --all-targets`
   
   A Criterion benchmark compares the direct parser with the previous 
Value-tree path. On this machine the direct path was faster for all covered 
inputs:
   
   | Input | Direct | Value tree | Improvement |
   |---|---:|---:|---:|
   | Integers | 534 ns | 644 ns | 17% |
   | Mixed object | 711 ns | 793 ns | 10% |
   | Fixed-point decimals | 437 ns | 474 ns | 8% |
   | Unescaped strings | 394 ns | 484 ns | 19% |
   | Escaped strings | 527 ns | 561 ns | 6% |
   | Nested object/list | 850 ns | 1,062 ns | 20% |
   
   The decimal baseline encodes non-integer numbers as `Double`, while the 
direct path performs the new exact decimal conversion.
   
   ## Are there any user-facing changes?
   
   Yes. Parsing fixed-point JSON text into Variant now preserves exact decimal 
semantics and precision instead of converting those values to `Double`.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to