adriangb opened a new pull request, #11236: URL: https://github.com/apache/arrow-rs/pull/11236
# Which issue does this PR close? - Part of https://github.com/apache/arrow-rs/issues/11234. # Rationale for this change A caller that reads ahead pushes the bytes of row groups before the decoder reaches them. If the caller then uses `into_builder` to skip some of those row groups (for example, after row-group pruning at runtime), the rebuilt decoder keeps their bytes until `clear_all_ranges` is called or the decoder is dropped. These bytes use memory and count against the caller's read-ahead budget, but the decoder never reads them. ```rust decoder.push_range(file_range, file_bytes)?; // prefetch let decoder = decoder .into_builder()? .with_row_groups(vec![0]) // skip row group 1 .build()?; decoder.buffered_bytes() // before: all prefetched bytes // after: only the column chunks of row group 0 ``` # What changes are included in this PR? | Change | Where | |---|---| | `PushBuffers::retain_ranges`: remove the buffered bytes outside the given ranges. The kept parts are zero-copy slices. | `util/push_buffers.rs` | | `ParquetPushDecoderBuilder::build` releases the buffered bytes that are outside the read column chunks (output and predicate columns) of the queued row groups. | `push_decoder/mod.rs`, `push_decoder/remaining.rs`, `push_decoder/reader_builder/mod.rs` | | Docs of `into_builder` and `with_buffers`. | `push_decoder/mod.rs` | This is useful for the default `RowGroup` granularity, and also for the batch granularity that https://github.com/apache/arrow-rs/issues/11234 adds. https://github.com/apache/arrow-rs/pull/11235 adds `PushBuffers::release_ranges` and the same `merge_ranges` helper in the same file. The PR that merges second needs a small rebase. # Are these changes tested? Yes. - `retain_ranges_keeps_only_the_given_bytes`: unit test. - `test_into_builder_releases_bytes_of_skipped_row_groups`: prefetch the file, rebuild to read only row group 0 with a wider projection, and check that only the column chunks of row group 0 stay buffered and that the decoder does not request them again. - `test_into_builder_preserves_buffered_bytes` now expects that the rebuild keeps the bytes of row group 1 and releases the rest of the prefetched file. Both decoder tests fail without the change. # Are there any user-facing changes? Yes, a behavior change: after `into_builder().build()`, the decoder no longer holds bytes that it does not read. The `into_builder` docs said that such bytes stay buffered until `clear_all_ranges`, and are updated. There is no API change. 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
