JingsongLi commented on PR #10173:
URL: https://github.com/apache/paimon/pull/10173#issuecomment-5952023758

   Reviewed head `24b212cfe1c45decf5225347ffff2905227d31af`. Honoring the 
existing table option has end-to-end value, and the earlier row-id-update 
finding is fixed. The native change is now correctly left to upstream #10202. 
However, enabling the existing full-stat collection/serialization helpers for 
`truncate(N)` exposes two unsafe persisted bounds that block production use.
   
   **[P1] Keep NaN in the newly emitted numeric bounds** 
(`data_writer.py:348-352`). With `metadata.stats-mode=truncate(3)`, a DOUBLE 
file containing `[1.0, NaN]` now records min=max `1.0`, because Arrow min/max 
ignore NaN; an all-NaN file records min=`+inf`, max=`-inf`. Java's comparison 
contract treats NaN above finite values. In an actual Python-written Avro 
table, Java's `equal(v, NaN)` scan/read returns the NaN row with `none` or 
`counts` (one candidate file), but `truncate(3)` prunes the entire file and 
returns no rows. The exact baseline wrote no value bounds for this option. 
Please compute compatible numeric bounds or conservatively omit them when NaN 
is present, including the all-NaN case.
   
   **[P1] Preserve nanosecond precision before publishing timestamp bounds** 
(same newly enabled write path). For TIMESTAMP(9) values epoch+1ns and 
epoch+2ns, the Parquet file retains both values, but the manifest min/max 
serializer records epoch for both. An actual `v > epoch` read on the exact 
baseline returns both rows; HEAD plans zero splits and returns nothing. Java 
independently reproduces the same result: baseline one file / IDs `[0,1]`, HEAD 
zero files / no IDs. Please preserve sub-microsecond precision or 
omit/conservatively widen bounds for these types before publishing them.
   
   The shared helpers already have these limitations under `full`; these 
findings concern this PR changing the existing `truncate(N)` option from safe 
no-bounds scans into incorrect pruning. They do not attribute the older 
full-mode behavior to this PR. I also observed a separate pre-existing Parquet 
NaN row-group filtering issue under `none`/`counts`; the NaN end-to-end witness 
above deliberately uses Avro to avoid relying on it.
   
   Validation: 95 Python tests passed (22 native-extension tests skipped 
locally); 12 Java collector tests passed on JDK 8 with normal Maven checks. 
Actual Java reads of four Python-written mode tables preserve 
string/binary/Unicode/null data, and full/truncate reduce an ordinary equality 
query from three candidate files to one. Independent probes passed 16 persisted 
PK/partial-update/blob/vector workflows across all modes, retaining exact key 
stats and correctly mapped value stats. 14,000 Unicode/binary truncation bound 
checks passed. Configured flake8, Python 3.6 grammar and diff checks passed; 
current-head CI is green. The cross-engine NaN and nanosecond regressions above 
are additional cases missing from those checks.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to