adriangb opened a new pull request, #11236:
URL: https://github.com/apache/arrow-rs/pull/11236

   # Which issue does this PR close?
   
   - Part of https://github.com/apache/arrow-rs/issues/11234.
   
   # Rationale for this change
   
   A caller that reads ahead pushes the bytes of row groups before the decoder 
reaches them. If the caller then uses `into_builder` to skip some of those row 
groups (for example, after row-group pruning at runtime), the rebuilt decoder 
keeps their bytes until `clear_all_ranges` is called or the decoder is dropped. 
These bytes use memory and count against the caller's read-ahead budget, but 
the decoder never reads them.
   
   ```rust
   decoder.push_range(file_range, file_bytes)?;     // prefetch
   let decoder = decoder
       .into_builder()?
       .with_row_groups(vec![0])                   // skip row group 1
       .build()?;
   decoder.buffered_bytes()
   // before: all prefetched bytes
   // after:  only the column chunks of row group 0
   ```
   
   # What changes are included in this PR?
   
   | Change | Where |
   |---|---|
   | `PushBuffers::retain_ranges`: remove the buffered bytes outside the given 
ranges. The kept parts are zero-copy slices. | `util/push_buffers.rs` |
   | `ParquetPushDecoderBuilder::build` releases the buffered bytes that are 
outside the read column chunks (output and predicate columns) of the queued row 
groups. | `push_decoder/mod.rs`, `push_decoder/remaining.rs`, 
`push_decoder/reader_builder/mod.rs` |
   | Docs of `into_builder` and `with_buffers`. | `push_decoder/mod.rs` |
   
   This is useful for the default `RowGroup` granularity, and also for the 
batch granularity that https://github.com/apache/arrow-rs/issues/11234 adds.
   
   https://github.com/apache/arrow-rs/pull/11235 adds 
`PushBuffers::release_ranges` and the same `merge_ranges` helper in the same 
file. The PR that merges second needs a small rebase.
   
   # Are these changes tested?
   
   Yes.
   
   - `retain_ranges_keeps_only_the_given_bytes`: unit test.
   - `test_into_builder_releases_bytes_of_skipped_row_groups`: prefetch the 
file, rebuild to read only row group 0 with a wider projection, and check that 
only the column chunks of row group 0 stay buffered and that the decoder does 
not request them again.
   - `test_into_builder_preserves_buffered_bytes` now expects that the rebuild 
keeps the bytes of row group 1 and releases the rest of the prefetched file.
   
   Both decoder tests fail without the change.
   
   # Are there any user-facing changes?
   
   Yes, a behavior change: after `into_builder().build()`, the decoder no 
longer holds bytes that it does not read. The `into_builder` docs said that 
such bytes stay buffered until `clear_all_ranges`, and are updated. There is no 
API change.
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to