rgbuilds opened a new pull request, #25822: URL: https://github.com/apache/datafusion/pull/25822
## Which issue does this PR close? - Closes #18355. ## Rationale for this change `row_groups_pruned_bloom_filter` currently reports retained row groups as Bloom filter matches even when Bloom pruning was not evaluated—for example, when Bloom filters are disabled or unavailable, no relevant Bloom statistics were loaded, or predicate evaluation failed. This makes `EXPLAIN ANALYZE` suggest that Bloom filters evaluated and matched row groups when they did not participate in the pruning decision. It also displays an idle Bloom pruning metric when all its counters are zero. A retained row group is not necessarily a Bloom filter match. The metric should describe actual Bloom pruning outcomes. ## What changes are included in this PR? - Record a Bloom filter match only when the Bloom pruning predicate is successfully evaluated and determines that the row group may match. - Leave Bloom pruning counters unchanged when Bloom pruning is unavailable or cannot be evaluated. - Continue retaining row groups conservatively when Bloom predicate evaluation fails, while recording the failure in `predicate_evaluation_errors`. - Omit `row_groups_pruned_bloom_filter` from physical-plan displays when all of its pruning counters are zero. - Preserve the registered metric and direct metric lookup; only its accounting and displayed output change. - Update affected Rust assertions and SQL logic test snapshots. The row-group access decisions and query results are unchanged. ## What is the testing strategy for this PR? Added `bloom_filter_pruning_error_retains_row_group_without_match`, which verifies that a Bloom predicate evaluation error: - retains the row group; - increments `predicate_evaluation_errors`; - does not increment either the Bloom matched or pruned counter. The display tests verify that: - an idle Bloom pruning metric is omitted from aggregated and full Indent, Graphviz, and PostgreSQL JSON output; - unrelated zero-valued metrics remain visible; - genuine nonzero Bloom pruning results remain visible. The following existing test coverage was also run: - Parquet Bloom-filter and row-group-filter unit tests; - `parquet_integration` row-group-pruning tests; - `core_integration` explain-analyze tests; - physical-plan display tests; - affected SQL logic tests for dynamic filtering, dynamic row-group pruning, explain-analyze, limit pruning, and Parquet filter pushdown. Formatting checks also pass. ## Are there any user-facing changes? Yes. Parquet Bloom filter pruning metrics in `EXPLAIN ANALYZE` and other physical-plan displays more accurately represent actual Bloom filter evaluation. An entirely idle `row_groups_pruned_bloom_filter` metric is no longer displayed. There are no public API changes and no changes to query results or row-group access decisions. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
