rich7420 opened a new pull request, #6397:
URL: https://github.com/apache/datafusion-comet/pull/6397

   ## Which issue does this PR close?
   
   Part of #6093. This independent optimization addresses an IF bottleneck 
encountered while validating #6180. It does not add AtLeastNNonNulls or depend 
on its implementation.
   
   ## Rationale for this change
   
   For mixed conditions, IF delegates to the pinned DataFusion CASE evaluator, 
which filters both branches and merges their results. Expressions such as 
`IF(isnan(c), 0D, c)` can select directly from a column and a literal instead. 
Repeating this expression across a wide aggregate exposed enough overhead to 
offset the benefit of a native filter in the #6180 integration workload.
   
   ## What changes are included in this PR?
   
   - Cache eligibility for one direct column and one non-NULL literal, in 
either branch order. Check exact types before evaluating the predicate so 
fallback cannot evaluate a nondeterministic predicate twice.
   - Return the selected scalar or column for uniform conditions. For mixed 
Float32/Float64/Int32/Int64 conditions, copy the value buffer once, replace 
literal-selected positions and update validity. Other eligible types use Arrow 
zip. NULL conditions select the false branch.
   - Keep computed/fallible branches, NULL literals, two columns/literals and 
coercion on CaseExpr with its existing short-circuit behavior.
   - Add SQL and Rust regression coverage plus a 160-case Criterion benchmark. 
No dependency upgrade, unsafe code or general CASE rewrite is included.
   
   ## How are these changes tested?
   
   [Fork CI at 
`21a44af8e`](https://github.com/rich7420/datafusion-comet/actions/runs/36542587652)
 passed native build, Rust tests, Spark 4.1 Comet suites, TPC-H/TPC-DS checks, 
lint and benchmark compilation. Spark's own 4.1 SQL suites passed Catalyst, all 
three SQL Core shards and all three Hive shards.
   
   Six focused release Rust tests passed locally, including sliced 
value/validity buffers, NULL conditions with true payload bits, dictionary 
logical NULLs, floating NaN payloads and signed-zero bits, integer limits, 
coercion and skipped/required ANSI cast errors. The combined #6180 candidate 
also passed 18 conditional SQL/AtLeastNNonNulls integration tests with 
native-path assertions. Workspace/all-target release Clippy, fmt, Spotless and 
RAT passed on the reviewed implementation.
   
   On Apple M3 Pro, for 8,192 Float64 rows with a random 5%-true mask, no NULLs 
and scalar/column branches, two paired release measurements improved from 
15.56–15.91 us to 2.51–2.57 us, about 6x. The benchmark includes small and full 
batches, uniform/mixed masks, nullable inputs and unchanged fallback controls. 
Two 2,950-case actual-library comparisons and two additional opaque-dispatch 
comparisons matched outputs exactly. Tiny/uniform controls retain nanosecond 
overhead; neither method repeated a greater-than-5% loss for a case whose 
baseline was at least 1 us. [Full Criterion tables and measurement 
qualifications](https://github.com/rich7420/datafusion-comet/pull/40) include 
slower cases as well as improvements.
   
   For context, the combined #6180 candidate on Spark 4.1/local[1] consumes 
count and 32 cleaned sums after `na.drop(16)` over 1,048,576 Parquet rows with 
32 Double columns. Two independent alternating-order runs measured 5% NaN 
native/fallback medians of 327/339 ms and 315/338 ms (3.6–7.0% faster); the 0% 
NaN controls also remained faster. Answers and native AtLeastNNonNulls/IF paths 
were checked. This measures the combined candidate against Comet fallback, not 
the standalone IF PR's end-to-end speedup.
   
   The branch is based on `ba9aa33db`. A merge-tree check against upstream main 
`8369bf11a` is conflict-free, with identical IF implementation/tests/benchmark 
and unchanged dependency pins. Main adds a separate benchmark registration 
elsewhere in Cargo.toml. Upstream CI will validate the merge result. Other 
Spark-profile SQL suites and Iceberg suites were not run for this candidate.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to