sunchao opened a new issue, #5507:
URL: https://github.com/apache/datafusion-comet/issues/5507

   ### Describe the bug
   
   Floating-point keys nested in arrays or structs can still produce different 
native `ORDER BY` and rank results from Spark. Spark treats signed zeros as 
equal and equates every NaN representation; Arrow's raw nested ordering can 
distinguish these representations. #5469 fixes scalar `FLOAT`/`DOUBLE` 
comparison keys, but its type gate deliberately does not recurse into arrays or 
structs.
   
   ### Steps to reproduce
   
   Use a Spark session with Comet enabled and `CometShuffleManager` configured. 
Create a Parquet input to preserve the zero signs and avoid folding the query 
to a local constant result:
   
   ```sql
   CREATE TABLE comet_nested_sort_case (id INT, v FLOAT) USING parquet;
   INSERT INTO comet_nested_sort_case VALUES
     (1, CAST('-0.0' AS FLOAT)), (2, CAST('0.0' AS FLOAT)), (3, 1.0);
   
   SET spark.sql.adaptive.enabled=false;
   SET spark.sql.shuffle.partitions=1;
   SET spark.comet.exec.strictFloatingPoint=false;
   
   SELECT id, array(v) AS k
   FROM comet_nested_sort_case
   ORDER BY k, id DESC;
   
   SELECT id, k, rnk FROM (
     SELECT id, array(v) AS k, RANK() OVER (ORDER BY array(v)) AS rnk
     FROM comet_nested_sort_case
   ) ranked WHERE rnk <= 1;
   ```
   
   Repeat both queries with `named_struct('value', v)` instead of `array(v)`, 
and with a `DOUBLE` input. Compare results with `spark.comet.enabled=false`. 
Inspect the executed plan and require `CometSortExec` in the native run; for 
the rank query, also record whether native 
`CometWindowExec`/`CometWindowGroupLimitExec` ran so a fallback does not hide 
the issue.
   
   As a safety control, repeat with `spark.comet.exec.strictFloatingPoint=true` 
and verify Spark fallback plus Spark-equivalent results. Keep the original zero 
signs in returned values. Additional regression data should construct distinct 
positive/negative NaN payloads programmatically; a plain SQL `'NaN'` literal 
does not exercise every representation.
   
   ### Expected behavior
   
   The two nested zero keys are peers. The `ORDER BY k, id DESC` result must 
put `id=2` before `id=1`, followed by `id=3`. The rank cutoff must retain both 
zero rows at rank 1. Equivalent expectations apply to `DENSE_RANK` and to 
compound nested keys.
   
   ### Additional context
   
   The current [key-construction 
gate](https://github.com/apache/datafusion-comet/blob/96eafdfe14384976cb18b7364c4cf84d7d1fb723/native/core/src/execution/planner.rs#L927-L963)
 only wraps scalar floats, and 
[normalize_array](https://github.com/apache/datafusion-comet/blob/96eafdfe14384976cb18b7364c4cf84d7d1fb723/native/spark-expr/src/math_funcs/internal/normalize_nan.rs#L57-L73)
 leaves other types unchanged. A fix must keep sorting, window peer comparison, 
rank-limit comparison, and relevant range-partition boundaries consistent while 
preserving original output values.
   
   This is separate from #5468's scalar fix. #5191 concerns 
`arrays_overlap`/`array_position`, not nested `ORDER BY` or window ranking. The 
conservative strict-mode fallback should remain until this ordering scope has 
its own correctness coverage. Fresh Spark 4.1.3 probes covered both FLOAT and 
DOUBLE: array/struct `ORDER BY` and array rank used native operators and 
mismatched Spark in six cases; struct rank already fell back and matched Spark 
for both widths. Strict-mode controls fell back and matched Spark in all eight 
query cases. Do not report that fallback case as a reproduced native wrong 
result.
   
   
   The runtime probe reused an existing OSS native build from #5420 with 
freshly compiled JVM code; it was not a full #5469 native rebuild. The 
inspected #5469 change only normalizes scalar keys, leaving these nested paths 
unchanged. This is a pre-existing limitation, not a regression introduced by 
the scalar fix.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to