tam3tamtam opened a new pull request, #51557:
URL: https://github.com/apache/arrow/pull/51557

   ### Rationale for this change
   
   Fixes #51548 
   
   The dataframe interchange protocol includes byte order in each dtype 
description. The PyArrow importer currently wraps these buffers directly, so 
non-native-endian values can be interpreted incorrectly and silently corrupted.
   
   ### What changes are included in this PR?
   
   The importer now converts non-native-endian data and string offset buffers 
to native byte order when needed, using `array.array`'s C-level `byteswap()`.
   If that conversion would require a copy and `allow_copy=False`, it raises a 
`RuntimeError`.
   
   Regression tests cover numeric values, string offsets, sentinel nulls, 16-bit
   and 64-bit values, and the `allow_copy=False` and invalid-buffer error paths.
   
   ### Are these changes tested?
   
   Yes. The focused tests for these changes passed:
   
   ```text
   python -m pytest pyarrow/tests/interchange/test_conversion.py -k 
"non_native_endian or unsupported_endianness or misaligned_buffer_size or 
int8_endianness" -q
   9 passed, 1792 deselected
   ```
   
   ### Are there any user-facing changes?
   
   Non-native-endian columns are now imported with their correct values. 
Importing
   them may copy their buffers; with `allow_copy=False`, conversion raises an 
error when a copy is required.
   
   This PR contains a "Critical Fix". Non-native-endian values could previously 
be interpreted with the wrong byte order and silently corrupted during 
dataframe interchange imports. This change preserves the producer's values.
   
   ### Was AI used for this PR?
   
   In accordance to the [AI generation 
guidelines](https://arrow.apache.org/docs/dev/developers/overview.html#ai-generated-code),
 please disclose below whether and how AI was used in this PR.
   
   **PR code and description written by:**
   
   - [x] Human
   - [x] AI
   
   **Reviewed before submission by:**
   
   - [x] Human
   - [x] AI
   - [ ] Not reviewed
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to