adriangb opened a new issue, #11234: URL: https://github.com/apache/arrow-rs/issues/11234
Part of https://github.com/apache/arrow-rs/issues/6946. # Goal Let `ParquetPushDecoder` request and decode one batch at a time, so that a caller gets the first batch of a row group as soon as the pages of that batch are pushed, and the decoder holds approximately one batch of bytes, not the full row group. ```rust let mut decoder = ParquetPushDecoderBuilder::try_new_decoder(metadata)? .with_fetch_granularity(FetchGranularity::Batch) // new, opt-in .build()?; // NeedsData asks only for the pages of the next batch. ``` Prototype measurement (one row group of 84 MB, 50 ms request latency): | Mode | First batch | Peak buffered bytes | |---|---:|---:| | `FetchGranularity::RowGroup` (default) | 451 ms | 84 MB | | `FetchGranularity::Batch`, read-ahead of 8 MB | 89 ms | 8.1 MB | https://github.com/apache/arrow-rs/pull/11223 has the full change (+4,156 lines). It is too large to review in one step, so this EPIC splits it. # Plan Each PR is useful without the PRs after it. No PR adds a temporary error or fallback that a later PR removes. A later PR can add an optimization to code that is already correct. | PR | Content | Base | Useful alone because | |---|---|---|---| | **A** https://github.com/apache/arrow-rs/pull/11218 | Test: the decoder needs the full row group before it returns a batch | `main` | Records the current behavior | | **C** TBD | `into_builder`: release the pushed bytes of row groups that the new decoder does not read | `main` | Today these bytes stay until `clear_all_ranges` | | **D** TBD | Keep `PushBuffers` sorted by offset | `main` | Lookups become a binary search. Today each lookup scans all buffers. | | **E** TBD | `FetchGranularity::Batch`, with and without a `RowFilter`. Holds the bytes of a row group until the row group ends. | A, #11233 fix | Lower time to the first batch | | **F** TBD | Batch mode: release each data page when all readers have passed it | E, D | Lower peak memory | | **G** TBD | Batch mode: use the predicate cache | E | Filtered scans do not decode predicate columns two times | | **H** TBD | Batch mode: use the configured `RowSelectionPolicy` (mask) in each window | E | Faster dense selections | ```text main ──┬── A test ─────────────────────┐ ├── #11233 fix ──────────────────┴── E batch granularity ──┬── F page release ◀── also needs D ├── C into_builder release ├── G predicate cache └── D sorted PushBuffers ──────────────────────────────────┴── H selection policy ``` A, C, D and the #11233 fix are independent and can be reviewed in parallel. F, G and H are independent of each other. # Tangential - https://github.com/apache/arrow-rs/issues/11233: the decoder does not release a pushed buffer that is larger than the requested ranges. This is a bug in the default mode. E uses the fix (`PushBuffers::release_ranges`), but the fix is not part of this EPIC. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
