Yahya Kisana created FLINK-40927:
------------------------------------

             Summary:  Inconsistent NaN semantics for FLOAT/DOUBLE across 
planner and runtime
                 Key: FLINK-40927
                 URL: https://issues.apache.org/jira/browse/FLINK-40927
             Project: Flink
          Issue Type: Bug
          Components: Table SQL / Planner
            Reporter: Yahya Kisana


Flink has no single definition of how NaN behaves for FLOAT/DOUBLE. Different 
parts of the planner and runtime treat it differently, so results can depend on 
the plan rather than the data.

||Area||NaN behaviour||Code||
|Comparisons in filters/projections|IEEE 754: {{NaN = NaN}} is 
FALSE|ScalarOperatorGens.scala:623-627|
|Planner simplification (Calcite)|Assumes values equal themselves and are 
totally ordered|RexSimplify.java (reflexive rewrite: FLINK-XXXXX; NOT/SEARCH 
negation: 1004-1019)|
|Sort comparator (ORDER BY, TopN)|NaN compares equal to every value (not a 
valid ordering)|GenerateUtils.scala:638-640|
|Batch sort normalized keys|NaN sorts above +Infinity|SortUtil.java:102-115|
|RecordEqualiser|Binary rows: NaN = NaN; other rows: NaN != 
NaN|EqualiserCodeGenerator.scala:72-74 vs 159-160|
|GROUP BY / DISTINCT / join keys|Byte comparison: NaNs with the same bits are 
equal|BinarySection.java:65-75|
|NaN constants|Cannot be represented 
(BigDecimal)|ValueLiteralExpression.java:216|



Other engines take one of two positions:
* IEEE 754 everywhere ({{NaN = NaN}} is FALSE): Trino and Impala (which 
disables affected simplifications).
* NaN is a normal value, equal to itself and greater than all other numbers: 
Spark/Databricks, PostgreSQL, DuckDB. This makes the planner's assumptions 
valid and gives sorting, grouping and joins one consistent order.

This was found while working on 
[FLINK-40923|https://issues.apache.org/jira/browse/FLINK-40923]

Needs further discussion



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to