JingsongLi commented on PR #10166: URL: https://github.com/apache/paimon/pull/10166#issuecomment-5831882161
Reviewed the current head as a data-correctness change, including the writer sort/fold path, read-side comparator, native dispatch boundaries, and the earlier review findings. There is clear end-to-end value: without this, buffering `(seq=100, high)` before `(seq=50, low)` can persist the wrong winner for a primary-key table. The current implementation orders buffered rows by PK, configured sequence fields, and generated sequence number; it uses explicit IEEE keys so NaNs and signed zero follow Java ordering. Invalid sequence configurations are checked before native selection and before dynamic-bucket row-key extraction. Floating native reads and floating/descending native writes fall back to the Python implementations where the native comparator is not equivalent. Local verification on this head: 183 focused sequence/native tests passed (14 optional-Rust cases skipped locally); 156 adjacent write-buffer, partial-update, aggregation, and table-write tests passed, plus 6 subtests. `git diff --check` passed. CI is green for Native CI and Python 3.6, 3.7, 3.10, 3.11, 3.12, and 3.13. The new real-table tests cover batch/stream and multiple write groupings, NaN/infinity/signed-zero/null ordering, compound fields, buffer flushes, and early invalid-configuration rejection. The three previously reported issues appear addressed at this head. I found no blocking regression. A Python-write/Java-read interoperability smoke test would add confidence in the shared-table contract, especially for floating sequences, but the current writer ordering and comparator agree with the Java path I inspected. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
