YinZheng-Sun opened a new pull request, #956:
URL: https://github.com/apache/iceberg-cpp/pull/956

   ## What
   
   Adds `ParquetMetricsRowGroupFilter` to evaluate Iceberg predicates against 
Parquet footer statistics before reading row groups.
   
   - Binds task filters against the complete table schema in 
`FileScanTaskReader`. Direct Parquet readers bind unbound filters against the 
projection in `Open()`.
   - Rewrites `NOT` predicates and evaluates comparisons, null checks, prefix 
predicates, and `IN` directly with Iceberg literals.
   - Applies the 200-element `IN` limit after binding and deduplication.
   - Reuses `ParquetMetrics::StatsValueToLiteral()` and 
`ValidateParquetTypeCompatibility()`.
   - Builds the field-ID-to-column-index mapping once per file and shares it 
across row-group evaluations.
   - Reads selected row groups individually, restoring each group's original 
starting row position.
   
   Adds tests covering split intersection, non-contiguous row groups, physical 
row positions, unprojected filter columns, schema evolution, missing 
statistics, predicate binding, and delete handling.
   
   ## Why
   
   The Parquet reader previously selected row groups only by file split. Filter 
predicates could not eliminate row groups whose footer statistics ruled out 
matching rows.
   
   Skipping groups also introduces gaps in physical row positions. Preserving 
those positions is necessary for `_pos`, inherited `_row_id`, and 
position-based delete handling.
   
   ## Behavior change
   
   - Enables statistics-based pruning by default through 
`read.parquet.row-group-filter.enabled`.
   - Retained row groups still require residual predicate evaluation; this does 
not implement exact row filtering.
   - Batches stop at row-group boundaries, including when pruning is disabled.
   - Unbound filter references missing from the projection fail during 
`Open()`. Callers must bind references to unprojected columns against the 
complete table schema.
   - Unsupported terms and unavailable statistics retain candidate groups 
conservatively.
   - Dictionary filtering, Bloom filtering, and page skipping remain deferred.
   
   Floating-point range pruning follows the examined iceberg-rust 
implementation. It retains a known correctness limitation: NaNs omitted from 
finite footer bounds can cause matching rows to be skipped. This case is not 
covered by the passing regression suite.
   
   ## Testing
   
   Built and ran `parquet_test`, `data_test`, `avro_test`, and 
`expression_test` locally during implementation. Rebuilt and reran 
`parquet_test` and `data_test` after the binding and field-mapping refactors; 
both passed.
   
   The `IN` regression covers 201 input literals that deduplicate to 200 
values, ensuring the limit applies to the bound set.
   
   Repository pre-commit checks, including clang-format and cmake-format, 
passed on the changed files. `git diff --check` is clean.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to