andygrove opened a new issue, #25913:
URL: https://github.com/apache/datafusion/issues/25913

   ### Describe the bug
   
   `datafusion-spark`'s `xxhash64` hashes a `Float32` or `Float64` NaN by its 
raw bits. Spark hashes floats through `Float.floatToIntBits` and 
`Double.doubleToLongBits`, which return the canonical NaN (`0x7fc00000`, 
`0x7ff8000000000000`) for every NaN, so in Spark all NaNs hash alike. A NaN 
with the sign bit set, or with another payload, therefore gets a different hash 
from `datafusion-spark` than from Spark.
   
   Such NaNs are common: negating a NaN sets its sign bit, and on x86-64 every 
NaN that arithmetic produces at run time, such as `0.0 / 0.0`, has the sign bit 
set.
   
   The float path is `hash_array_primitive_float!` in 
`datafusion/spark/src/function/hash/utils.rs`. It hashes `0` for `-0.0`, as 
Spark does, and otherwise the bytes of the value as they are.
   
   ### To Reproduce
   
   ```rust
   let values: ArrayRef = Arc::new(Float64Array::from(vec![
       f64::NAN,                              // 0x7ff8000000000000, the 
canonical NaN
       f64::from_bits(0xfff8_0000_0000_0000), // the same with the sign bit 
set, which -f64::NAN gives
       f64::from_bits(0x7ff0_0000_0000_0001), // a NaN with another payload
   ]));
   let args = ScalarFunctionArgs {
       args: vec![ColumnarValue::Array(Arc::clone(&values))],
       arg_fields: vec![Arc::new(Field::new("v", DataType::Float64, true))],
       number_rows: values.len(),
       return_field: Arc::new(Field::new("r", DataType::Int64, false)),
       config_options: Arc::new(ConfigOptions::default()),
   };
   let hashes = SparkXxhash64::new()
       .invoke_with_args(args)?
       .into_array(values.len())?;
   ```
   
   `hashes` is `[-3127944061524951246, 9200374361256412029, 
5729085064965309005]`. `Float32` behaves the same way: `f32::NAN` hashes to 
`2692338816207849720` and the NaN with the sign bit set (`0xffc00000`) to 
`4760557555880201639`.
   
   ### Expected behavior
   
   Every NaN hashes like the canonical NaN. In Spark, when `d` is NaN, 
`xxhash64(d)` and `xxhash64(-d)` both return `-3127944061524951246`.
   
   Canonicalizing NaN next to the existing `-0.0` check in 
`hash_array_primitive_float!` would match Spark.
   
   ### Additional context
   
   Comet found this while fixing the same gap in its own kernel 
(apache/datafusion-comet#6413), which keeps float arguments off `SparkXxhash64` 
until it is fixed.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to