rich7420 opened a new pull request, #6396:
URL: https://github.com/apache/datafusion-comet/pull/6396

   ## Which issue does this PR close?
   
   Part of #6093. Extracts the generic filter output optimization from #6180, 
following [review 
feedback](https://github.com/apache/datafusion-comet/pull/6180#discussion_r4127884161).
 This PR does not close the native AtLeastNNonNulls issue.
   
   ## Rationale for this change
   
   A native filter currently materializes every input column even when its 
parent only consumes a subset, or only needs the row count. Wide inputs with 
narrow projections spend time filtering arrays whose values are never read.
   
   ## What changes are included in this PR?
   
   Push distinct required column indices from a column-only projection into 
DataFusion's filter output projection. Keep the parent projection to restore 
aliases, order and duplicate outputs. The predicate retains its original input 
schema, and both native plans remain under the existing Spark metrics nodes, 
including shared plan IDs. Empty projections preserve row counts. Computed 
projections and projections requiring every input column retain their existing 
behavior.
   
   ## How are these changes tested?
   
   - [Fork CI at 
`256239ebe`](https://github.com/rich7420/datafusion-comet/actions/runs/36542579994)
 passed native build, Rust tests, Spark 4.1 Comet suites, TPC-H/TPC-DS result 
checks and lint. Spark's own 4.1 SQL suites also passed: Catalyst, all three 
SQL Core shards and all three Hive shards.
   - Locally, 50 Rust planner tests and both focused Spark SQL/metrics tests 
passed. Planner assertions check the actual filter projection, output mapping, 
zero-column row counts and shared/distinct plan IDs. SQL cases cover 
reordered/duplicate outputs, aliases, NULLs, empty results and a rejected row 
containing an invalid ANSI cast. The Scala test checks native filter/project 
metrics.
   - Prior measurements of this pruning implementation reduced narrow-output 
string query time by about 19–22%, with full-output controls unchanged. Those 
measurements used the #6180 workload; they are not fresh measurements of this 
extracted branch.
   
   The branch is based on `ba9aa33db`. A merge-tree check against upstream main 
`8369bf11a` is conflict-free and retains the same implementation and test 
files. New main changes overlap only in a different section of the operator 
guide. Upstream CI will validate that merge result. Other Spark-profile SQL 
suites and Iceberg suites were not run for this candidate.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to