peterxcli opened a new issue, #11260:
URL: https://github.com/apache/arrow-rs/issues/11260

   ### Is your feature request related to a problem or challenge?
   
   Comet currently calls `unshred_variant`, then decodes and rebuilds its 
output to match Spark's Variant bytes (apache/datafusion-comet#5978). We need 
to reuse Arrow's typed decoders and recursive traversal while writing the final 
representation directly.
   
   The public API materializes a VariantArray using Arrow's writer. The 
reusable 
[`UnshredVariantRowBuilder`](https://github.com/apache/arrow-rs/blob/76829d5ce4562616a69ea8b2696e96160514482f/parquet-variant-compute/src/unshred_variant.rs#L160-L207)
 is private. Making it public alone would not provide enough control: typed and 
residual scalars share `append_value`, and nested output uses concrete 
object/list builders.
   
   ### Describe the solution you'd like
   
   Expose a reusable row decoder in `parquet-variant-compute`, constructed once 
per input array, with a fallible row visitor or equivalent borrowed interface. 
For example, `decode_row(index, visitor)`; the exact API is open for discussion.
   
   The caller needs to:
   
   - Receive typed scalars separately from residual bytes and their source 
metadata. Spark narrows typed integers and canonicalizes typed NaNs, while 
preserving those encodings in residual scalars.
   - Handle nested object fields and list elements through the same interface, 
preserving typed schema order and original residual field order.
   - Distinguish a null parent, a missing shredding state and explicit Variant 
null, with access to state validity before residual decoding. Honor parent 
masking and sliced list offsets.
   - Build output metadata independently of input metadata. Residual containers 
must expose their names/children so the caller can remap field IDs instead of 
copying containers with stale IDs.
   
   Keep `unshred_variant` as the existing canonical entry point, ideally backed 
by this decoder and its default writer. Spark-specific encoding and 
permissive-reading policies remain explicit consumer choices.
   
   An external-crate test should demonstrate direct output for partially and 
fully shredded nested rows, without an intermediate encoded VariantArray. Cover 
typed/residual provenance, differing dictionaries, nulls, sliced lists and 
fallible malformed-input handling; verify the existing entry point retains its 
behavior. A focused benchmark should report allocation and runtime differences.
   
   ### Describe alternatives you've considered
   
   Unshred then re-encode, as Comet does today, or duplicate Arrow's typed 
decoders downstream. Both add avoidable work or maintenance.
   
   #10620 addresses returning a materialized composite Variant from 
`try_value`; this request exposes traversal before materialization and could 
share its decoder.
   
   ### Additional context
   
   apache/datafusion-comet#5978 tracks the full Spark-compatible reconstruction 
and performance target. This API is one prerequisite; legacy input handling and 
other compatibility policies still need separate integration.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to