rajaryan2007 opened a new pull request, #51647:
URL: https://github.com/apache/arrow/pull/51647
### Rationale for this change
Casting a `list_view` to a `list` (or across offset widths) incorrectly
reused the offsets buffer and ignored the `sizes` buffer. This caused
out-of-bounds memory reads (crashes) and generated invalid data whenever views
were overlapping or out-of-order.
```python
views = pa.ListViewArray.from_arrays(
pa.array([2, 0, 1], pa.int32()), # offsets
pa.array([2, 2, 1], pa.int32()), # sizes
pa.array([1, 2, 3, 4], pa.int32()))
views.cast(pa.list_(pa.int32())).to_pylist()
# Before: [[], [1], []] (Invalid: negative / non-monotonic offsets)
# After: [[3, 4], [1, 2], [2]] (Correct)
What changes are included in this PR?
list_util::internal::ListFromListView: New shared internal helper that
materializes a standard list layout by gathering the values referenced by each
view. Supports all 4 offset-width combinations.
CastList Update: List-view sources now use this helper before casting the
child array.
FromListView Fix: ListArray::FromListView and LargeListArray::FromListView
now delegate to this shared helper. This removes duplicated code (-31 lines)
and fixes a bug where nulls were handled incorrectly for sliced inputs
(GH-51613).
Are these changes tested?
Yes. Added and updated tests in C++ (scalar_cast_test.cc, list_util_test.cc,
array_list_test.cc) and Python (test_array.py). Tests cover all offset-width
combinations, null views, out-of-order/overlapping views, sliced inputs,
zero-length edge cases, and nested list-views.
Are there any user-facing changes?
Yes.
Performance: list_view -> list casts are no longer zero-copy. The referenced
values are now always gathered into a new child array. This is a necessary
performance cost to guarantee correctness.
This PR includes breaking changes to public APIs.
ListArray::FromListView and LargeListViewArray::FromListView now return
correct nulls for sliced inputs. This changes the output for users who may have
been relying on the previous bugged behavior.
This PR contains a "Critical Fix".
Yes, this fixes a bug that produced incorrect and invalid data
(negative/non-monotonic offsets that fail validation) and prevents Check
failed: (off) <= (length) abort crashes caused by out-of-bounds reads.
Was AI used for this PR?
I used AI to help explain concepts, generate some inline code comments, and
assist with writing and building tests.
PR code and description written by:
[x] Human
[ ] AI
Reviewed before submission by:
[x] Human
[ ] AI
[ ] Not reviewed
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]