sunchao commented on code in PR #5457:
URL: https://github.com/apache/datafusion-comet/pull/5457#discussion_r4189010570
##########
native/core/src/parquet/schema_adapter.rs:
##########
@@ -1457,6 +1457,10 @@ impl SparkPhysicalExprAdapter {
| (DataType::Map(_, _), DataType::Map(_, _))
| (DataType::Timestamp(_, _), DataType::Timestamp(_, _))
| (DataType::Timestamp(_, _), DataType::Int64)
+ | (
+ DataType::Date32,
Review Comment:
Fixed in 633a805889. The full Spark 4.1 query reproduced the reported
NULL-versus-overflow mismatch. Filtered scans requesting NTZ fields now retain
Spark’s reader, including nested structs, arrays, and map keys/values. This
preserves selected-row overflow errors without eagerly failing on rows Spark
skips.
The conservative limitation is explicit in the code, compatibility guide,
and PR description: physical file types are unavailable at planning, so genuine
NTZ files also fall back when filtered. No planning-time footer IO is added.
Unfiltered DATE adaptation/NTZ scans and projections that omit NTZ remain
native.
The expanded regression passes with Parquet filter pushdown enabled and
disabled: both overflow signs raise plain ArithmeticException("long overflow"),
a filtered LIMIT skips the late overflow, genuine nested NTZ filtering falls
back, and the native controls remain native. Strict Spark 3.5 reactor test
compilation, Spotless and Scalastyle also pass. Native code is unchanged by
this fix; the tests reused the matching PR-built JNI library. Broader hosted
checks are pending.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]