YinZheng-Sun opened a new pull request, #956: URL: https://github.com/apache/iceberg-cpp/pull/956
## What Adds `ParquetMetricsRowGroupFilter` to evaluate Iceberg predicates against Parquet footer statistics before reading row groups. - Binds task filters against the complete table schema in `FileScanTaskReader`. Direct Parquet readers bind unbound filters against the projection in `Open()`. - Rewrites `NOT` predicates and evaluates comparisons, null checks, prefix predicates, and `IN` directly with Iceberg literals. - Applies the 200-element `IN` limit after binding and deduplication. - Reuses `ParquetMetrics::StatsValueToLiteral()` and `ValidateParquetTypeCompatibility()`. - Builds the field-ID-to-column-index mapping once per file and shares it across row-group evaluations. - Reads selected row groups individually, restoring each group's original starting row position. Adds tests covering split intersection, non-contiguous row groups, physical row positions, unprojected filter columns, schema evolution, missing statistics, predicate binding, and delete handling. ## Why The Parquet reader previously selected row groups only by file split. Filter predicates could not eliminate row groups whose footer statistics ruled out matching rows. Skipping groups also introduces gaps in physical row positions. Preserving those positions is necessary for `_pos`, inherited `_row_id`, and position-based delete handling. ## Behavior change - Enables statistics-based pruning by default through `read.parquet.row-group-filter.enabled`. - Retained row groups still require residual predicate evaluation; this does not implement exact row filtering. - Batches stop at row-group boundaries, including when pruning is disabled. - Unbound filter references missing from the projection fail during `Open()`. Callers must bind references to unprojected columns against the complete table schema. - Unsupported terms and unavailable statistics retain candidate groups conservatively. - Dictionary filtering, Bloom filtering, and page skipping remain deferred. Floating-point range pruning follows the examined iceberg-rust implementation. It retains a known correctness limitation: NaNs omitted from finite footer bounds can cause matching rows to be skipped. This case is not covered by the passing regression suite. ## Testing Built and ran `parquet_test`, `data_test`, `avro_test`, and `expression_test` locally during implementation. Rebuilt and reran `parquet_test` and `data_test` after the binding and field-mapping refactors; both passed. The `IN` regression covers 201 input literals that deduplicate to 200 values, ensuring the limit applies to the bound set. Repository pre-commit checks, including clang-format and cmake-format, passed on the changed files. `git diff --check` is clean. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
