neilconway opened a new pull request, #11323:
URL: https://github.com/apache/arrow-rs/pull/11323

   # Which issue does this PR close?
   
   - N/A
   
   # Rationale for this change
   
   Arrow does not require bytes outside the visible span of the values buffer 
to be valid UTF-8. This rule was not consistently applied:
   
   * `ArrayData` checked for UTF-8 validity over the visible span, but 
   * `StringArray::try_new` and `try_from_binary` checked for UTF-8 validity of 
the entire values buffer, which is both inefficient for sliced arrays and also 
rejects otherwise valid inputs
   * `substring` assumed that the entire values buffer contained valid UTF-8 
bytes, which was sound for inputs from the `StringArray` path but not for the 
`ArrayData` path
   * Binary to Utf8 casts (in `safe: true`) reserved enough space for the whole 
values buffer, not the visible span
   
   # What changes are included in this PR?
   
   * Both validators now check only the bytes between the first and last offset
   * `substring` reads only those bytes as a `&str`.
   * The `safe: true` fallback of the Binary to Utf8 cast also reserves only 
the bytes the offsets span, not the whole values buffer.
   * Add unit tests
   
   # Are these changes tested?
   
   Yes; existing tests pass, new tests added.
   
   # Are there any user-facing changes?
   
   No.
   
   # AI usage
   
   Developed with Claude Code (Opus 5.5), reviewed with Codex (Astra 6). I 
revised and understand the resulting code.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to