rich7420 opened a new pull request, #6397: URL: https://github.com/apache/datafusion-comet/pull/6397
## Which issue does this PR close? Part of #6093. This independent optimization addresses an IF bottleneck encountered while validating #6180. It does not add AtLeastNNonNulls or depend on its implementation. ## Rationale for this change For mixed conditions, IF delegates to the pinned DataFusion CASE evaluator, which filters both branches and merges their results. Expressions such as `IF(isnan(c), 0D, c)` can select directly from a column and a literal instead. Repeating this expression across a wide aggregate exposed enough overhead to offset the benefit of a native filter in the #6180 integration workload. ## What changes are included in this PR? - Cache eligibility for one direct column and one non-NULL literal, in either branch order. Check exact types before evaluating the predicate so fallback cannot evaluate a nondeterministic predicate twice. - Return the selected scalar or column for uniform conditions. For mixed Float32/Float64/Int32/Int64 conditions, copy the value buffer once, replace literal-selected positions and update validity. Other eligible types use Arrow zip. NULL conditions select the false branch. - Keep computed/fallible branches, NULL literals, two columns/literals and coercion on CaseExpr with its existing short-circuit behavior. - Add SQL and Rust regression coverage plus a 160-case Criterion benchmark. No dependency upgrade, unsafe code or general CASE rewrite is included. ## How are these changes tested? [Fork CI at `21a44af8e`](https://github.com/rich7420/datafusion-comet/actions/runs/36542587652) passed native build, Rust tests, Spark 4.1 Comet suites, TPC-H/TPC-DS checks, lint and benchmark compilation. Spark's own 4.1 SQL suites passed Catalyst, all three SQL Core shards and all three Hive shards. Six focused release Rust tests passed locally, including sliced value/validity buffers, NULL conditions with true payload bits, dictionary logical NULLs, floating NaN payloads and signed-zero bits, integer limits, coercion and skipped/required ANSI cast errors. The combined #6180 candidate also passed 18 conditional SQL/AtLeastNNonNulls integration tests with native-path assertions. Workspace/all-target release Clippy, fmt, Spotless and RAT passed on the reviewed implementation. On Apple M3 Pro, for 8,192 Float64 rows with a random 5%-true mask, no NULLs and scalar/column branches, two paired release measurements improved from 15.56–15.91 us to 2.51–2.57 us, about 6x. The benchmark includes small and full batches, uniform/mixed masks, nullable inputs and unchanged fallback controls. Two 2,950-case actual-library comparisons and two additional opaque-dispatch comparisons matched outputs exactly. Tiny/uniform controls retain nanosecond overhead; neither method repeated a greater-than-5% loss for a case whose baseline was at least 1 us. [Full Criterion tables and measurement qualifications](https://github.com/rich7420/datafusion-comet/pull/40) include slower cases as well as improvements. For context, the combined #6180 candidate on Spark 4.1/local[1] consumes count and 32 cleaned sums after `na.drop(16)` over 1,048,576 Parquet rows with 32 Double columns. Two independent alternating-order runs measured 5% NaN native/fallback medians of 327/339 ms and 315/338 ms (3.6–7.0% faster); the 0% NaN controls also remained faster. Answers and native AtLeastNNonNulls/IF paths were checked. This measures the combined candidate against Comet fallback, not the standalone IF PR's end-to-end speedup. The branch is based on `ba9aa33db`. A merge-tree check against upstream main `8369bf11a` is conflict-free, with identical IF implementation/tests/benchmark and unchanged dependency pins. Main adds a separate benchmark registration elsewhere in Cargo.toml. Upstream CI will validate the merge result. Other Spark-profile SQL suites and Iceberg suites were not run for this candidate. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
