Cintu07 opened a new pull request, #11259: URL: https://github.com/apache/arrow-rs/pull/11259
### Which issue does this PR close? - Follow-up to #11224, part of #11011. Nothing to close here. ### Rationale for this change a view addresses its value with a block id and a 32 bit offset into that buffer. the row decoder wrote every non inlined value into one buffer and hardcoded the block id to 0, so once those values passed i32::MAX bytes there was no offset left that could reach them. #11224 turns that into an error. view arrays carry a list of data buffers for exactly this case, so those rows are decodable, they just need the next buffer. ### What changes are included in this PR? when the value just appended would not fit the current buffer, it moves to a fresh one and the block id goes up with it. one compare per row, one memcpy per rolled value. each buffer asks for min(long bytes left, i32::MAX) rather than the whole total, so the allocation stays inside a buffer as well. the error stays for a single value longer than one whole buffer, since one view cannot span two of them. stacked on #11224, so the first commit here is theirs. ### Are these changes tested? yes, two tests in arrow-row. decode_binary_view_capped takes the buffer length so a test can force a roll without decoding two gigabytes, the same way gc_with_max_buffer_size is tested in arrow-array. the roundtrip covers ascending, descending and both null orders, and asserts the uncapped path still gives one buffer. putting the block id back to 0 makes it fail. ### Are there any user-facing changes? no new api. #11224 already changed decode_binary_view and decode_string_view to return Result, this only widens what decodes without an error. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
