rajaryan2007 opened a new pull request, #51647:
URL: https://github.com/apache/arrow/pull/51647

   ### Rationale for this change
   
   Casting a `list_view` to a `list` (or across offset widths) incorrectly 
reused the offsets buffer and ignored the `sizes` buffer. This caused 
out-of-bounds memory reads (crashes) and generated invalid data whenever views 
were overlapping or out-of-order. 
   
   ```python
   views = pa.ListViewArray.from_arrays(
       pa.array([2, 0, 1], pa.int32()),   # offsets
       pa.array([2, 2, 1], pa.int32()),   # sizes
       pa.array([1, 2, 3, 4], pa.int32()))
   views.cast(pa.list_(pa.int32())).to_pylist()
   
   # Before: [[], [1], []]         (Invalid: negative / non-monotonic offsets)
   # After:  [[3, 4], [1, 2], [2]] (Correct)
   
   What changes are included in this PR?
   list_util::internal::ListFromListView: New shared internal helper that 
materializes a standard list layout by gathering the values referenced by each 
view. Supports all 4 offset-width combinations.
   
   CastList Update: List-view sources now use this helper before casting the 
child array.
   
   FromListView Fix: ListArray::FromListView and LargeListArray::FromListView 
now delegate to this shared helper. This removes duplicated code (-31 lines) 
and fixes a bug where nulls were handled incorrectly for sliced inputs 
(GH-51613).
   
   Are these changes tested?
   Yes. Added and updated tests in C++ (scalar_cast_test.cc, list_util_test.cc, 
array_list_test.cc) and Python (test_array.py). Tests cover all offset-width 
combinations, null views, out-of-order/overlapping views, sliced inputs, 
zero-length edge cases, and nested list-views.
   
   Are there any user-facing changes?
   Yes.
   
   Performance: list_view -> list casts are no longer zero-copy. The referenced 
values are now always gathered into a new child array. This is a necessary 
performance cost to guarantee correctness.
   
   This PR includes breaking changes to public APIs.
   ListArray::FromListView and LargeListViewArray::FromListView now return 
correct nulls for sliced inputs. This changes the output for users who may have 
been relying on the previous bugged behavior.
   
   This PR contains a "Critical Fix".
   Yes, this fixes a bug that produced incorrect and invalid data 
(negative/non-monotonic offsets that fail validation) and prevents Check 
failed: (off) <= (length) abort crashes caused by out-of-bounds reads.
   
   Was AI used for this PR?
   I used AI to help explain concepts, generate some inline code comments, and 
assist with writing and building tests.
   
   PR code and description written by:
   
   [x] Human
   [ ] AI
   
   Reviewed before submission by:
   
   [x] Human
   [ ] AI
   [ ] Not reviewed


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to