pingzh opened a new pull request, #6067:
URL: https://github.com/apache/datafusion-comet/pull/6067

   ## Which issue does this PR close?
   
   Part 2 of the [five-PR plan for 
#5775](https://github.com/apache/datafusion-comet/issues/5775#issuecomment-5672549320),
 following the shared runtime-filter refactor in #5937. #5775 remains open for 
TopK fusion and reader pushdown.
   
   ## Rationale for this change
   
   Verification on Apache `main` at `6065705c16340c0be293212a71decfd9df4daae4` 
reproduced error suppression in existing join runtime filtering.
   
   A Spark 4.1 broadcast join reads a Parquet INT64 payload as INT32 and has no 
matching probe keys. Vanilla Spark and Comet with runtime filtering disabled 
report `SchemaColumnConvertNotSupportedException`. Enabling the join runtime 
filter incorrectly lets the query succeed because reader pruning skips the 
incompatible data.
   
   Native join tests also reproduce suppressed type-promotion errors in 
mixed-schema files and suppressed millisecond-to-microsecond timestamp 
overflow, including nested timestamps. These tests fail on the unmodified 
production code.
   
   ## What changes are included in this PR?
   
   - Wrap the existing schema-adapter factory for scans with an attached 
runtime filter. Check the projected columns and existing predicate dependencies 
separately for each file.
   - Keep reader filtering for direct column mappings and literals. For other 
adaptations, replace the reader's dynamic predicate with `true` and retain the 
original adapter, static predicates, and normal conversion-error timing. 
Decoded-batch runtime filtering still applies.
   - Fall back to decoded-batch filtering when supplied file statistics can 
prune before the per-file adapter runs.
   - Add native and Spark regressions and document the conservative fallback. 
No new configuration or metrics.
   
   Empty files, statically excluded files, unprojected mismatches, and allowed 
conversions remain readable. Existing direct-schema reader-pruning tests still 
pass.
   
   The separate missing-null-statistics investigation passed all 48 join 
configurations (96 executions). The fix therefore does not add the 
footer/statistics rewriting from the earlier TopK implementation.
   
   ## How are these changes tested?
   
   - Before the fix: four native regression tests failed on unchanged `main`; 
the readable controls and statistics investigation passed. The Spark regression 
failed specifically with Comet runtime filtering enabled after verifying that 
vanilla Spark and Comet filtering disabled raised the expected error.
   - After the fix: all 24 focused native tests passed, covering existing 
runtime-filter behavior, the new regressions, and planner metric export.
   - All 14 targeted Spark join dynamic-filter tests passed with Spark 4.1.3 
and JDK 21, including the regression above. Verified that the packaged JNI 
library matches the rebuilt native library.
   - Core crate Clippy passed for all targets with warnings denied; Cargo 
formatting, Maven Spotless/Scalastyle, Prettier, and `git diff --check` passed.
   - Three independent sub-agents reviewed the native guard; the Spark 
regression and documentation also received independent review.
   - Applied `run-spark-4.1-tests` to request the broader Spark SQL suite 
required for native-operator changes; its CI verdict is pending.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to