spancer opened a new issue, #10334: URL: https://github.com/apache/paimon/issues/10334
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon/issues) and found nothing similar. ### Paimon version 2.0.0 (paimon-bundle-2.0.0) ### Compute Engine None — reproduced with the Java API, local filesystem catalog, Java 21. ### Minimal reproduce step Write the same two rows to an ORC table and a Parquet table and compare file statistics and query results: ```java for (String format : new String[] {"orc", "parquet"}) { Identifier id = Identifier.create("db", "t_nan_" + format); catalog.createTable( id, Schema.newBuilder() .column("id", DataTypes.INT()) .column("c", DataTypes.DOUBLE()) .option("file.format", format) .build(), false); Table table = catalog.getTable(id); write(table, List.of(GenericRow.of(1, 1.0), GenericRow.of(2, Double.NaN))); for (Split split : table.newReadBuilder().newScan().plan().splits()) { for (DataFileMeta file : ((DataSplit) split).dataFiles()) { System.out.println( format + " min=" + file.valueStats().minValues().getDouble(1) + " max=" + file.valueStats().maxValues().getDouble(1)); } } PredicateBuilder builder = new PredicateBuilder(table.rowType()); System.out.println(format + " c > 2.0 : " + count(table, builder.greaterThan(1, 2.0))); System.out.println(format + " c = NaN : " + count(table, builder.equal(1, Double.NaN))); } ``` Output: ```text orc min=1.0 max=1.0 orc c > 2.0 : 0 orc c = NaN : 0 parquet min=1.0 max=NaN parquet c > 2.0 : 2 parquet c = NaN : 2 ``` The same happens with signed zero. Writing `(0.0, -0.0)` produces: ```text orc min=0.0 max=0.0 -> c < 0.0 returns 0 rows parquet min=-0.0 max=0.0 -> c < 0.0 returns 2 rows (file not skipped) ``` Helper methods: ```java static void write(Table table, List<GenericRow> rows) throws Exception { BatchWriteBuilder builder = table.newBatchWriteBuilder(); try (BatchTableWrite write = builder.newWrite(); BatchTableCommit commit = builder.newCommit()) { for (GenericRow row : rows) write.write(row); commit.commit(write.prepareCommit()); } } static long count(Table table, Predicate filter) throws Exception { ReadBuilder builder = table.newReadBuilder(); if (filter != null) builder = builder.withFilter(filter); long[] count = new long[1]; try (RecordReader<InternalRow> reader = builder.newRead().createReader(builder.newScan().plan())) { reader.forEachRemaining(row -> count[0]++); } return count[0]; } ``` ### What doesn't meet your expectations? On ORC the file-level statistics are not bounds under the ordering Paimon uses to evaluate predicates on DOUBLE / FLOAT (`Double.compareTo`: NaN is greater than every other value, and -0.0 is less than 0.0): - with `(1.0, NaN)` the max is recorded as `1.0`, so the file is skipped for `c > 2.0` and `c = NaN`, although the NaN row satisfies both; - with `(0.0, -0.0)` the min is recorded as `0.0`, so the file is skipped for `c < 0.0`, although the -0.0 row satisfies it. The Parquet path records `max = NaN` and `min = -0.0` for the same data and does not skip the file (it returns the whole file, which is fine for a pushed-down filter). The same table definition therefore gives different query results depending only on `file.format`. ### Anything else? `org.apache.paimon.format.orc.filter.OrcSimpleStatsExtractor` passes ORC's own double column statistics through unchanged. ORC updates them with primitive `<` / `>` comparisons, so a NaN never becomes min or max unless it is the first value, and -0.0 never replaces 0.0. Suggested fix: make the ORC statistics consistent with the Parquet path, or, when that is not possible, do not emit min/max for FLOAT / DOUBLE columns on ORC. A missing bound is safe; a wrong bound silently drops rows. ### Are you willing to submit a PR? - [ ] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
