Yahya Kisana created FLINK-40927:
------------------------------------
Summary: Inconsistent NaN semantics for FLOAT/DOUBLE across
planner and runtime
Key: FLINK-40927
URL: https://issues.apache.org/jira/browse/FLINK-40927
Project: Flink
Issue Type: Bug
Components: Table SQL / Planner
Reporter: Yahya Kisana
Flink has no single definition of how NaN behaves for FLOAT/DOUBLE. Different
parts of the planner and runtime treat it differently, so results can depend on
the plan rather than the data.
||Area||NaN behaviour||Code||
|Comparisons in filters/projections|IEEE 754: {{NaN = NaN}} is
FALSE|ScalarOperatorGens.scala:623-627|
|Planner simplification (Calcite)|Assumes values equal themselves and are
totally ordered|RexSimplify.java (reflexive rewrite: FLINK-XXXXX; NOT/SEARCH
negation: 1004-1019)|
|Sort comparator (ORDER BY, TopN)|NaN compares equal to every value (not a
valid ordering)|GenerateUtils.scala:638-640|
|Batch sort normalized keys|NaN sorts above +Infinity|SortUtil.java:102-115|
|RecordEqualiser|Binary rows: NaN = NaN; other rows: NaN !=
NaN|EqualiserCodeGenerator.scala:72-74 vs 159-160|
|GROUP BY / DISTINCT / join keys|Byte comparison: NaNs with the same bits are
equal|BinarySection.java:65-75|
|NaN constants|Cannot be represented
(BigDecimal)|ValueLiteralExpression.java:216|
Other engines take one of two positions:
* IEEE 754 everywhere ({{NaN = NaN}} is FALSE): Trino and Impala (which
disables affected simplifications).
* NaN is a normal value, equal to itself and greater than all other numbers:
Spark/Databricks, PostgreSQL, DuckDB. This makes the planner's assumptions
valid and gives sorting, grouping and joins one consistent order.
This was found while working on
[FLINK-40923|https://issues.apache.org/jira/browse/FLINK-40923]
Needs further discussion
--
This message was sent by Atlassian Jira
(v8.20.10#820010)