mbutrovich opened a new pull request, #3302: URL: https://github.com/apache/iceberg-rust/pull/3302
## Which issue does this PR close? - Closes #3298. ## What changes are included in this PR? - Add `residual_for_missing_fields` in `arrow/reader/predicate_visitor.rs`. It replaces each predicate leaf on a top-level field that is missing from the data file with `AlwaysTrue` or `AlwaysFalse`, evaluated against the value projection returns for that field: the identity partition value, otherwise `initial-default`. Leaves on missing fields with neither keep the existing null handling. This is the same idea as Java's `ResidualEvaluator`, extended to `initial-default`. - Apply the residual in the reader pipeline right after the field ID map is built, so the row filter, row group pruning, bloom filter pruning, and page index pruning all see the same predicate. This includes the predicate built from equality deletes. - Leave leaves on nested fields unchanged, because a nested field also reads as null in any row where an ancestor struct is null. This keeps filtering consistent with projection, which doesn't apply nested defaults yet (#3261). - Reuse `ExpressionEvaluatorVisitor` for leaf evaluation and `constants_map` for identity partition values. Both, and `LogicalExpression::new`, go from private to `pub(crate)`. There are no public API changes. ## Are these changes tested? Yes, with unit tests and a `MemoryCatalog` scan test: - The three reproduction tests from #3298 in `row_filter.rs` cover the row filter with `initial-default`, page index pruning with `initial-default`, and the row filter with an identity partition value. The page index test now applies the residual before calling `get_row_selection_for_filter_predicate`, as the pipeline does. - Residual unit tests in `predicate_visitor.rs` follow Java's [`TestMetricsRowGroupFilter`](https://github.com/apache/iceberg/blob/48330b8dacab6662242d252b39c8444190979bb2/data/src/test/java/org/apache/iceberg/data/TestMetricsRowGroupFilter.java#L475-L641) cases for string, date, and double defaults. Other cases check that the partition value takes precedence over `initial-default`, that bucket partition values are ignored, and that a present column, a column with no value, and a nested field are left unchanged. They also check that `NOT`, `AND`, and `OR` keep their structure. - `test_scan_filter_on_initial_default_column_absent_from_file` in `scan/mod.rs` follows Java's [`testFilterPushdownOnInitialDefaultColumnAbsentFromFile`](https://github.com/apache/iceberg/blob/48330b8dacab6662242d252b39c8444190979bb2/spark/v4.1/spark/src/test/java/org/apache/iceberg/spark/sql/TestFilterPushDown.java#L651-L691). It appends a file, adds a column with a default through `update_schema`, appends a second file, and scans with filters, with row selection off and on. It fails without the fix. ## AI Disclosure I wrote this PR with help from Claude but understand and support the core approach, implementation, and test coverage. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
