CurtHagenlocher opened a new pull request, #450:
URL: https://github.com/apache/arrow-dotnet/pull/450

   ## What's Changed
   
   `ArrayDataConcatenator` (and so `ArrowArrayConcatenator`) couldn't handle 
two array types:
   
   - **Null arrays** had no visitor case and threw `NotImplementedException`, 
at the top level or as a child (e.g. a null-typed struct field, which Parquet 
readers hit for pyarrow `pa.null()` columns). They now concatenate to a null 
array of the combined length.
   - **Dictionary arrays** went through `Visit(FixedWidthType)` because 
`DictionaryType` derives from `FixedWidthType`. That concatenated the indices 
but dropped the dictionary, so `ArrowArrayConcatenator` threw `Dictionary must 
not be null` and `ArrayDataConcatenator` silently returned dictionary-typed 
data with no dictionary.
   
   For dictionaries:
   - If every non-empty input shares the same dictionary (the same `ArrayData`, 
or distinct `ArrayData` over the same memory, e.g. after 
`Retain`/`SliceShared`), the indices are concatenated and that dictionary is 
kept.
   - Otherwise the dictionaries of the non-empty inputs are concatenated and 
each input's indices are shifted by the combined length of the dictionaries 
before it. Null slots get index 0 instead of a shifted, possibly out-of-range 
value. If the combined dictionary is too big for the index type, this throws 
`OverflowException` instead of wrapping around.
   - Inputs with different index types or value types are rejected with 
`ArgumentException`.
   
   Dictionaries that are equal in content but stored separately take the 
concatenation path. The result is correct but not deduplicated. Unifying 
dictionaries could be added later if needed.
   
   New tests cover null arrays, a struct with a null field, shared and 
different dictionaries (including nulls, slices and empty inputs), index 
overflow at the `int8` boundary, and `ArrayDataConcatenator` keeping the 
dictionary.
   
   Closes #446.
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to