JingsongLi commented on PR #10173: URL: https://github.com/apache/paimon/pull/10173#issuecomment-5952023758
Reviewed head `24b212cfe1c45decf5225347ffff2905227d31af`. Honoring the existing table option has end-to-end value, and the earlier row-id-update finding is fixed. The native change is now correctly left to upstream #10202. However, enabling the existing full-stat collection/serialization helpers for `truncate(N)` exposes two unsafe persisted bounds that block production use. **[P1] Keep NaN in the newly emitted numeric bounds** (`data_writer.py:348-352`). With `metadata.stats-mode=truncate(3)`, a DOUBLE file containing `[1.0, NaN]` now records min=max `1.0`, because Arrow min/max ignore NaN; an all-NaN file records min=`+inf`, max=`-inf`. Java's comparison contract treats NaN above finite values. In an actual Python-written Avro table, Java's `equal(v, NaN)` scan/read returns the NaN row with `none` or `counts` (one candidate file), but `truncate(3)` prunes the entire file and returns no rows. The exact baseline wrote no value bounds for this option. Please compute compatible numeric bounds or conservatively omit them when NaN is present, including the all-NaN case. **[P1] Preserve nanosecond precision before publishing timestamp bounds** (same newly enabled write path). For TIMESTAMP(9) values epoch+1ns and epoch+2ns, the Parquet file retains both values, but the manifest min/max serializer records epoch for both. An actual `v > epoch` read on the exact baseline returns both rows; HEAD plans zero splits and returns nothing. Java independently reproduces the same result: baseline one file / IDs `[0,1]`, HEAD zero files / no IDs. Please preserve sub-microsecond precision or omit/conservatively widen bounds for these types before publishing them. The shared helpers already have these limitations under `full`; these findings concern this PR changing the existing `truncate(N)` option from safe no-bounds scans into incorrect pruning. They do not attribute the older full-mode behavior to this PR. I also observed a separate pre-existing Parquet NaN row-group filtering issue under `none`/`counts`; the NaN end-to-end witness above deliberately uses Avro to avoid relying on it. Validation: 95 Python tests passed (22 native-extension tests skipped locally); 12 Java collector tests passed on JDK 8 with normal Maven checks. Actual Java reads of four Python-written mode tables preserve string/binary/Unicode/null data, and full/truncate reduce an ordinary equality query from three candidate files to one. Independent probes passed 16 persisted PK/partial-update/blob/vector workflows across all modes, retaining exact key stats and correctly mapped value stats. 14,000 Unicode/binary truncation bound checks passed. Configured flake8, Python 3.6 grammar and diff checks passed; current-head CI is green. The cross-engine NaN and nanosecond regressions above are additional cases missing from those checks. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
