adriangb opened a new pull request, #11240:
URL: https://github.com/apache/arrow-rs/pull/11240

   > [!NOTE]
   > This PR is stacked on https://github.com/apache/arrow-rs/pull/11238. Only 
the last commit is new (+101/−9). I will rebase when 
https://github.com/apache/arrow-rs/pull/11238 merges.
   
   # Which issue does this PR close?
   
   - Part of https://github.com/apache/arrow-rs/issues/11234 (step G).
   
   # Rationale for this change
   
   With `FetchGranularity::Batch` from 
https://github.com/apache/arrow-rs/pull/11238, the output reader decodes the 
predicate columns that it reads again, because the batch-granular decoder does 
not use the predicate cache. The row-group mode uses the cache. With this PR, 
both modes use it.
   
   # What changes are included in this PR?
   
   | Change | Why |
   |---|---|
   | Each predicate reader writes the columns of `cache_projection` to a 
`RowGroupCache`. The output reader reads them from it. | The same cache setup 
as the row-group mode (`compute_cache_projection`, `max_predicate_cache_size`) |
   | The fetch of a cached column uses cache batch boundaries (`fetch_ranges` 
with the cache projection) | On a cache miss, the output reads the column again 
for the full cache batch. The decoder must hold those pages. |
   | Windows start at a multiple of `batch_size` | Already true. It aligns 
windows with the cache batches. Now documented. |
   
   The design is documented in *Predicate cache* in the module documentation of 
`reader_builder/incremental.rs`.
   
   https://github.com/apache/arrow-rs/pull/11239 (page release) must release a 
cached column at cache-batch boundaries. The PR that merges second adds this 
rule to `release_passed_pages`.
   
   # Are these changes tested?
   
   Yes.
   
   - `predicate_cache_is_used`: the output reads rows from the cache 
(`records_read_from_cache > 0`). It fails if the readers do not use the cache.
   - `predicate_cache_misses`: a cache of 0 and 1 bytes forces misses. The 
batches are the same as with `RowGroup`. This test, and 3 of the equivalence 
tests, fail if the fetch does not use cache batch boundaries.
   
   # Are there any user-facing changes?
   
   No API change. Filtered scans with `FetchGranularity::Batch` decode less.
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to