alamb opened a new issue, #11029:
URL: https://github.com/apache/arrow-rs/issues/11029

   **Is your feature request related to a problem or challenge? Please describe 
what you are trying to do.**
   
   https://github.com/apache/arrow-rs/pull/10842 introduces the 
[`PageIndexProvider`](https://github.com/apache/arrow-rs/blob/79dbf5a3524cb8077d58eaec21c17f97de4ba4ca/parquet/src/file/metadata/page_index.rs)
 trait so that applications can plug in custom page index implementations (for 
example, sparse or lazily-loaded indexes) into 
[`ParquetMetaData`](https://github.com/apache/arrow-rs/blob/79dbf5a3524cb8077d58eaec21c17f97de4ba4ca/parquet/src/file/metadata/mod.rs).
   
   However, no test currently drives reading with a custom 
[`PageIndexProvider`](https://github.com/apache/arrow-rs/blob/79dbf5a3524cb8077d58eaec21c17f97de4ba4ca/parquet/src/file/metadata/page_index.rs)
 through the actual read path.
   
   Several places in the code use `as_any().downcast_ref::<PageIndex>()` and 
take a different path for custom providers (for example 
[`ParquetMetaData::memory_size`](https://github.com/apache/arrow-rs/blob/79dbf5a3524cb8077d58eaec21c17f97de4ba4ca/parquet/src/file/metadata/mod.rs),
 `PartialEq for ParquetMetaData`, and 
[`ParquetMetaDataWriter`](https://github.com/apache/arrow-rs/blob/79dbf5a3524cb8077d58eaec21c17f97de4ba4ca/parquet/src/file/metadata/writer.rs)),
 so adding specific reader coverage would be good.
   
   - The 
[`custom_page_index.rs`](https://github.com/apache/arrow-rs/blob/79dbf5a3524cb8077d58eaec21c17f97de4ba4ca/parquet/examples/custom_page_index.rs)
 example implements `OnDemandPageIndexProvider`, but it only prints index 
availability, and examples are compiled — not executed — in CI.
   - The existing reader tests exercise the built-in `PageIndex` implementation 
only.
   
   
   **Describe the solution you'd like**
   
   Add an integration test that:
   
   1. Writes a parquet file with page indexes (multiple row groups, multiple 
columns, multiple pages per column).
   2. Loads the metadata with 
[`PageIndexPolicy::Skip`](https://github.com/apache/arrow-rs/blob/79dbf5a3524cb8077d58eaec21c17f97de4ba4ca/parquet/src/file/metadata/reader.rs),
 then installs a custom `PageIndexProvider` implementation (similar to the 
example's 
[`OnDemandPageIndexProvider`](https://github.com/apache/arrow-rs/blob/79dbf5a3524cb8077d58eaec21c17f97de4ba4ca/parquet/examples/custom_page_index.rs))
 via 
[`ParquetMetaDataBuilder::set_page_index`](https://github.com/apache/arrow-rs/blob/79dbf5a3524cb8077d58eaec21c17f97de4ba4ca/parquet/src/file/metadata/mod.rs),
 populated for only a subset of row groups / columns.
   3. Reads data through 
[`ParquetRecordBatchReader`](https://github.com/apache/arrow-rs/blob/79dbf5a3524cb8077d58eaec21c17f97de4ba4ca/parquet/src/arrow/arrow_reader/mod.rs)
 (and/or the async reader / push decoder) with a 
[`RowSelection`](https://github.com/apache/arrow-rs/blob/79dbf5a3524cb8077d58eaec21c17f97de4ba4ca/parquet/src/arrow/arrow_reader/selection.rs)
 that triggers page skipping, and verifies:
      - correct results are returned,
      - the offset-index-driven fetch path is used for columns where the 
provider returns `Some`,
      - columns where the provider returns `None` fall back to 
whole-column-chunk fetching.
   
   **Describe alternatives you've considered**
   
   Converting the 
[`custom_page_index.rs`](https://github.com/apache/arrow-rs/blob/79dbf5a3524cb8077d58eaec21c17f97de4ba4ca/parquet/examples/custom_page_index.rs)
 example into an executed test, or adding assertions to it and running it in CI.
   
   **Additional context**
   
   - Follow-up from review of https://github.com/apache/arrow-rs/pull/10842
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to