andygrove opened a new issue, #6702: URL: https://github.com/apache/datafusion-comet/issues/6702
### What is the problem the feature request solves? With `spark.comet.parquet.rowFilterPushdown.enabled=true`, the native Parquet reader drops the rows that the scan's data filters reject, so a data filter that compares a `FLOAT` or `DOUBLE` column has to follow Spark's semantics on every row. With #6447, the planner normalizes both operands of such a comparison in that mode (`FloatNormalize [child: d@0] > 500`). A raw column would drop rows that Spark matches, such as a stored NaN with the sign bit set, which Arrow orders below every other value. DataFusion's `PruningPredicate` can't see through the normalization, so for these comparisons it is `always_true()`, and they get no row group statistics, page index or bloom filter pruning. With the flag off, which is the default, the data filters only prune, so they compare the raw column with the constant and keep pruning, and the Filter above the scan applies Spark's semantics. In DataFusion 55, one predicate drives both pruning and the row filter in `ParquetSource`, so keeping pruning in this mode needs a pruning predicate that is separate from the row filter. ### Describe the potential solution A DataFusion change that lets `ParquetSource` take a pruning predicate alongside the row filter predicate. Comet would pass the raw comparison for pruning and the normalized one for filtering rows. Part of this could also be done in Comet alone. `=`, `<>`, `<=>` and `IS DISTINCT FROM` against a constant other than NaN could keep the raw column with row-level pushdown too: DataFusion compares `-0.0` and `0.0` as equal, and no NaN equals such a constant, so these comparisons already keep the rows Spark keeps. That brings back pruning for equality, but not for `<`, `<=`, `>` and `>=`. ### Additional context Came up in the review of #6447: https://github.com/apache/datafusion-comet/pull/6447#discussion_r4190049382. The floating-point compatibility guide describes the limitation and links here. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
