spancer opened a new issue, #10334:
URL: https://github.com/apache/paimon/issues/10334

   ### Search before asking
   
   - [x] I searched in the [issues](https://github.com/apache/paimon/issues) 
and found nothing similar.
   
   
   ### Paimon version
   
   2.0.0 (paimon-bundle-2.0.0)
   
   ### Compute Engine
   
   None — reproduced with the Java API, local filesystem catalog, Java 21.
   
   ### Minimal reproduce step
   
   Write the same two rows to an ORC table and a Parquet table and compare file 
statistics and query results:
   
   ```java
   for (String format : new String[] {"orc", "parquet"}) {
     Identifier id = Identifier.create("db", "t_nan_" + format);
     catalog.createTable(
         id,
         Schema.newBuilder()
             .column("id", DataTypes.INT())
             .column("c", DataTypes.DOUBLE())
             .option("file.format", format)
             .build(),
         false);
     Table table = catalog.getTable(id);
     write(table, List.of(GenericRow.of(1, 1.0), GenericRow.of(2, Double.NaN)));
   
     for (Split split : table.newReadBuilder().newScan().plan().splits()) {
       for (DataFileMeta file : ((DataSplit) split).dataFiles()) {
         System.out.println(
             format
                 + " min=" + file.valueStats().minValues().getDouble(1)
                 + " max=" + file.valueStats().maxValues().getDouble(1));
       }
     }
     PredicateBuilder builder = new PredicateBuilder(table.rowType());
     System.out.println(format + " c > 2.0 : " + count(table, 
builder.greaterThan(1, 2.0)));
     System.out.println(format + " c = NaN : " + count(table, builder.equal(1, 
Double.NaN)));
   }
   ```
   
   Output:
   
   ```text
   orc min=1.0 max=1.0
   orc c > 2.0 : 0
   orc c = NaN : 0
   parquet min=1.0 max=NaN
   parquet c > 2.0 : 2
   parquet c = NaN : 2
   ```
   
   The same happens with signed zero. Writing `(0.0, -0.0)` produces:
   
   ```text
   orc min=0.0 max=0.0          -> c < 0.0 returns 0 rows
   parquet min=-0.0 max=0.0     -> c < 0.0 returns 2 rows (file not skipped)
   ```
   
   Helper methods:
   
   ```java
   static void write(Table table, List<GenericRow> rows) throws Exception {
     BatchWriteBuilder builder = table.newBatchWriteBuilder();
     try (BatchTableWrite write = builder.newWrite();
         BatchTableCommit commit = builder.newCommit()) {
       for (GenericRow row : rows) write.write(row);
       commit.commit(write.prepareCommit());
     }
   }
   
   static long count(Table table, Predicate filter) throws Exception {
     ReadBuilder builder = table.newReadBuilder();
     if (filter != null) builder = builder.withFilter(filter);
     long[] count = new long[1];
     try (RecordReader<InternalRow> reader =
         builder.newRead().createReader(builder.newScan().plan())) {
       reader.forEachRemaining(row -> count[0]++);
     }
     return count[0];
   }
   ```
   
   ### What doesn't meet your expectations?
   
   On ORC the file-level statistics are not bounds under the ordering Paimon 
uses to evaluate predicates on DOUBLE / FLOAT (`Double.compareTo`: NaN is 
greater than every other value, and -0.0 is less than 0.0):
   
   - with `(1.0, NaN)` the max is recorded as `1.0`, so the file is skipped for 
`c > 2.0` and `c = NaN`, although the NaN row satisfies both;
   - with `(0.0, -0.0)` the min is recorded as `0.0`, so the file is skipped 
for `c < 0.0`, although the -0.0 row satisfies it.
   
   The Parquet path records `max = NaN` and `min = -0.0` for the same data and 
does not skip the file (it returns the whole file, which is fine for a 
pushed-down filter). The same table definition therefore gives different query 
results depending only on `file.format`.
   
   ### Anything else?
   
   `org.apache.paimon.format.orc.filter.OrcSimpleStatsExtractor` passes ORC's 
own double column statistics through unchanged. ORC updates them with primitive 
`<` / `>` comparisons, so a NaN never becomes min or max unless it is the first 
value, and -0.0 never replaces 0.0.
   
   Suggested fix: make the ORC statistics consistent with the Parquet path, or, 
when that is not possible, do not emit min/max for FLOAT / DOUBLE columns on 
ORC. A missing bound is safe; a wrong bound silently drops rows.
   
   ### Are you willing to submit a PR?
   
   - [ ] I'm willing to submit a PR!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to