[ 
https://issues.apache.org/jira/browse/FLINK-40927?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Yahya Kisana updated FLINK-40927:
---------------------------------
    Description: 
Flink has no single definition of how NaN behaves for FLOAT/DOUBLE. Different 
parts of the planner and runtime treat it differently, so results can depend on 
the plan rather than the data.

||Area||NaN behaviour||Code||
|Comparisons in filters/projections|IEEE 754: {{NaN = NaN}} is 
FALSE|[ScalarOperatorGens.scala:623-627|https://github.com/apache/flink/blob/master/flink-table/flink-table-planner/src/main/scala/org/apache/flink/table/planner/codegen/calls/ScalarOperatorGens.scala#L623-L627]|
|Planner simplification (Calcite)|Assumes values equal themselves and are 
totally ordered|RexSimplify.java (reflexive rewrite: FLINK-XXXXX; NOT/SEARCH 
negation: 1004-1019)|
|Sort comparator (ORDER BY, TopN)|NaN compares equal to every value (not a 
valid ordering)|GenerateUtils.scala:638-640|
|Batch sort normalized keys|NaN sorts above +Infinity|SortUtil.java:102-115|
|RecordEqualiser|Binary rows: NaN = NaN; other rows: NaN != 
NaN|EqualiserCodeGenerator.scala:72-74 vs 159-160|
|GROUP BY / DISTINCT / join keys|Byte comparison: NaNs with the same bits are 
equal|BinarySection.java:65-75|
|NaN constants|Cannot be represented 
(BigDecimal)|ValueLiteralExpression.java:216|



Other engines take one of two positions:
* IEEE 754 everywhere ({{NaN = NaN}} is FALSE): Trino and Impala (which 
disables affected simplifications).
* NaN is a normal value, equal to itself and greater than all other numbers: 
Spark/Databricks, PostgreSQL, DuckDB. This makes the planner's assumptions 
valid and gives sorting, grouping and joins one consistent order.

This was found while working on 
[FLINK-40923|https://issues.apache.org/jira/browse/FLINK-40923]

Needs further discussion

  was:
Flink has no single definition of how NaN behaves for FLOAT/DOUBLE. Different 
parts of the planner and runtime treat it differently, so results can depend on 
the plan rather than the data.

||Area||NaN behaviour||Code||
|Comparisons in filters/projections|IEEE 754: {{NaN = NaN}} is 
FALSE|ScalarOperatorGens.scala:623-627|
|Planner simplification (Calcite)|Assumes values equal themselves and are 
totally ordered|RexSimplify.java (reflexive rewrite: FLINK-XXXXX; NOT/SEARCH 
negation: 1004-1019)|
|Sort comparator (ORDER BY, TopN)|NaN compares equal to every value (not a 
valid ordering)|GenerateUtils.scala:638-640|
|Batch sort normalized keys|NaN sorts above +Infinity|SortUtil.java:102-115|
|RecordEqualiser|Binary rows: NaN = NaN; other rows: NaN != 
NaN|EqualiserCodeGenerator.scala:72-74 vs 159-160|
|GROUP BY / DISTINCT / join keys|Byte comparison: NaNs with the same bits are 
equal|BinarySection.java:65-75|
|NaN constants|Cannot be represented 
(BigDecimal)|ValueLiteralExpression.java:216|



Other engines take one of two positions:
* IEEE 754 everywhere ({{NaN = NaN}} is FALSE): Trino and Impala (which 
disables affected simplifications).
* NaN is a normal value, equal to itself and greater than all other numbers: 
Spark/Databricks, PostgreSQL, DuckDB. This makes the planner's assumptions 
valid and gives sorting, grouping and joins one consistent order.

This was found while working on 
[FLINK-40923|https://issues.apache.org/jira/browse/FLINK-40923]

Needs further discussion


>  Inconsistent NaN semantics for FLOAT/DOUBLE across planner and runtime
> -----------------------------------------------------------------------
>
>                 Key: FLINK-40927
>                 URL: https://issues.apache.org/jira/browse/FLINK-40927
>             Project: Flink
>          Issue Type: Bug
>          Components: Table SQL / Planner
>            Reporter: Yahya Kisana
>            Priority: Major
>
> Flink has no single definition of how NaN behaves for FLOAT/DOUBLE. Different 
> parts of the planner and runtime treat it differently, so results can depend 
> on the plan rather than the data.
> ||Area||NaN behaviour||Code||
> |Comparisons in filters/projections|IEEE 754: {{NaN = NaN}} is 
> FALSE|[ScalarOperatorGens.scala:623-627|https://github.com/apache/flink/blob/master/flink-table/flink-table-planner/src/main/scala/org/apache/flink/table/planner/codegen/calls/ScalarOperatorGens.scala#L623-L627]|
> |Planner simplification (Calcite)|Assumes values equal themselves and are 
> totally ordered|RexSimplify.java (reflexive rewrite: FLINK-XXXXX; NOT/SEARCH 
> negation: 1004-1019)|
> |Sort comparator (ORDER BY, TopN)|NaN compares equal to every value (not a 
> valid ordering)|GenerateUtils.scala:638-640|
> |Batch sort normalized keys|NaN sorts above +Infinity|SortUtil.java:102-115|
> |RecordEqualiser|Binary rows: NaN = NaN; other rows: NaN != 
> NaN|EqualiserCodeGenerator.scala:72-74 vs 159-160|
> |GROUP BY / DISTINCT / join keys|Byte comparison: NaNs with the same bits are 
> equal|BinarySection.java:65-75|
> |NaN constants|Cannot be represented 
> (BigDecimal)|ValueLiteralExpression.java:216|
> Other engines take one of two positions:
> * IEEE 754 everywhere ({{NaN = NaN}} is FALSE): Trino and Impala (which 
> disables affected simplifications).
> * NaN is a normal value, equal to itself and greater than all other numbers: 
> Spark/Databricks, PostgreSQL, DuckDB. This makes the planner's assumptions 
> valid and gives sorting, grouping and joins one consistent order.
> This was found while working on 
> [FLINK-40923|https://issues.apache.org/jira/browse/FLINK-40923]
> Needs further discussion



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to