CurtHagenlocher opened a new pull request, #447:
URL: https://github.com/apache/arrow-dotnet/pull/447

   ## What's Changed
   
   The shredding APIs only worked one value at a time, so every caller holding 
a `VariantArray` had to write the same loops: decode each row (plus a null 
mask) to shred it, and rebuild row by row to unshred it. This adds extension 
methods on `VariantArray` in `VariantArrayShreddingExtensions`:
   
   - `Reassemble(allocator = null)` converts a shredded array into its 
unshredded equivalent. Null elements stay null, and an unshredded input is 
returned unchanged.
   - `Shred(schema, allocator = null)` shreds into the given layout. It reads 
logical values, so an already-shredded input is reshredded rather than losing 
its typed columns.
   - `TryShred(options, out shredded, allocator = null)` infers a schema and 
shreds into it. It returns false (with `shredded` set to null) when the 
inferred schema is unshredded.
   - `InferShredSchema(options = null)` infers a schema from an array. This 
isn't in the issue's list, but it's needed for the batch scenario the issue 
describes: a Parquet writer infers once over a representative batch, then calls 
`Shred(schema)` on every batch so all row groups share one layout.
   
   All four go through one shared loop that resolves the column's schema and 
child arrays once rather than per row. Null elements pass through the 
`VariantValue?` pipeline added in #445.
   
   Closes #399.
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to