pingzh opened a new issue, #6741:
URL: https://github.com/apache/datafusion-comet/issues/6741

   ### Describe the bug
   
   Native `make_date` has two Spark compatibility gaps on `main` at 
`8d0bb01a5a99924024e07c1722e6eddd9260e167`:
   
   1. A valid date accepted by Java `LocalDate.of` whose epoch day exceeds 
Date32 is treated as an invalid calendar date. Comet returns NULL with ANSI 
disabled, or raises a date-field error with ANSI enabled. Spark raises 
`ARITHMETIC_OVERFLOW` in both modes.
   2. Native scalar-function argument evaluation does not preserve Spark's year 
→ month → day NULL checks. A computed month/day can throw even when Spark skips 
it after an earlier NULL argument. A projection below LIMIT can also evaluate a 
throwing row beyond the requested limit within an already-polled batch.
   
   The range-conversion behavior was reproduced by executing the exact upstream 
date helpers. The NULL/LIMIT findings are based on source inspection; the 
complete Apache Spark/JNI SQL reproductions below have not yet been run.
   
   ### Steps to reproduce
   
   Run with Comet configured and native execution enabled. Verify that the 
physical plan contains a native Comet projection for `make_date`; Spark 
fallback or constant folding can otherwise hide the divergence. Use Parquet 
columns to keep the constructor from folding during optimization.
   
   **Date32 overflow:**
   
   ```sql
   CREATE TABLE make_date_overflow_repro(y INT, m INT, d INT) USING parquet;
   INSERT INTO make_date_overflow_repro VALUES (6000000, 1, 1);
   
   SET spark.sql.ansi.enabled=false;
   SELECT make_date(y, m, d) FROM make_date_overflow_repro;
   
   SET spark.sql.ansi.enabled=true;
   SELECT make_date(y, m, d) FROM make_date_overflow_repro;
   ```
   
   For `(6000000, 1, 1)`, the upstream calendar helper returns epoch day 
`2190735472`, larger than `2147483647`. Native `make_date` narrows this and 
returns `None`. The caller maps this to NULL with ANSI disabled and to 
`DatetimeFieldOutOfBounds` with the incorrect detail `Invalid date 'JANUARY 1'` 
with ANSI enabled. On Spark 4.x, the latter becomes 
`DATETIME_FIELD_OUT_OF_BOUNDS`, rather than `ARITHMETIC_OVERFLOW`.
   
   **NULL-skipped argument divergence:**
   
   ```sql
   CREATE TABLE make_date_null_repro(y INT, m INT, inner_y INT) USING parquet;
   INSERT INTO make_date_null_repro VALUES (NULL, 1, 2024), (2024, NULL, 2024);
   
   SET spark.sql.ansi.enabled=true;
   SELECT make_date(y, m, dayofmonth(make_date(inner_y, 13, 1)))
   FROM make_date_null_repro;
   ```
   
   Spark should return two NULL rows without evaluating the invalid inner date. 
Native eager argument evaluation can instead raise its invalid-month error.
   
   **LIMIT evaluation-order check:**
   
   Construct a physical `Limit(Projection(input))` with a single batch 
containing a safe first row followed by a row whose `make_date` throws. Request 
only the first row and compare with Spark's bounded row consumption. Native 
projection currently evaluates the entire batch before the limit slices it. 
This depends on physical plan placement and batch boundaries; it does not imply 
that later, unrequested batches are always evaluated.
   
   ### Expected behavior
   
   - Dates accepted by Java `LocalDate.of` but outside Date32 raise Spark's 
`ARITHMETIC_OVERFLOW` in both ANSI modes.
   - Invalid calendar fields retain Spark's existing ANSI-dependent NULL/error 
behavior.
   - A NULL year skips month/day evaluation; a NULL month skips day evaluation. 
Computed arguments must preserve those masks, or the affected expression should 
retain Spark evaluation.
   - Lazy consumers preserve Spark's demanded-row behavior. Correcting overflow 
must not introduce errors from arguments or rows Spark skips.
   
   ### Additional context
   
   The earlier wide-year fix #5443 addressed Spark-valid dates inside Date32, 
including year 300000. Its [review explicitly deferred the beyond-Date32 
mismatch](https://github.com/apache/datafusion-comet/pull/5443#pullrequestreview-5002872377).
   
   Public source evidence:
   
   - [Date32 narrowing and NULL/error 
handling](https://github.com/apache/datafusion-comet/blob/8d0bb01a5a99924024e07c1722e6eddd9260e167/native/spark-expr/src/datetime_funcs/make_date.rs#L99-L182):
 `i32::try_from(days).ok()` loses the overflow distinction.
   - [CometMakeDate 
serialization](https://github.com/apache/datafusion-comet/blob/8d0bb01a5a99924024e07c1722e6eddd9260e167/spark/src/main/scala/org/apache/comet/serde/datetime.scala#L438-L454)
 and [native 
planning](https://github.com/apache/datafusion-comet/blob/8d0bb01a5a99924024e07c1722e6eddd9260e167/native/core/src/execution/planner.rs#L3957-L3963):
 all children become ordinary scalar-function arguments.
   - [DataFusion argument 
evaluation](https://github.com/apache/datafusion/blob/55.1.0/datafusion/physical-expr/src/scalar_function.rs#L236-L261):
 arguments are evaluated before invoking the UDF and before Comet checks their 
NULLs.
   - [Spark ternary NULL 
checks](https://github.com/apache/spark/blob/v4.0.1/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/Expression.scala#L848-L860)
 and [generated 
checks](https://github.com/apache/spark/blob/v4.0.1/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/Expression.scala#L895-L915).
   - [Spark 
MakeDate](https://github.com/apache/spark/blob/c7d67e3f5d4c9d88a480367b44fc54d26adf99ab/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/datetimeExpressions.scala#L2525-L2532)
 catches date-field exceptions; 
[localDateToDays](https://github.com/apache/spark/blob/c7d67e3f5d4c9d88a480367b44fc54d26adf99ab/sql/api/src/main/scala/org/apache/spark/sql/catalyst/util/SparkDateTimeUtils.scala#L163)
 uses exact integer narrowing, so arithmetic overflow propagates in either ANSI 
mode.
   - [DataFusion 
projection](https://github.com/apache/datafusion/blob/55.1.0/datafusion/physical-plan/src/projection.rs#L679-L703)
 evaluates a whole batch; 
[limit](https://github.com/apache/datafusion/blob/55.1.0/datafusion/physical-plan/src/limit.rs#L650-L693)
 slices afterward.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to