JingsongLi opened a new pull request, #1051: URL: https://github.com/apache/paimon-rust/pull/1051
### Purpose Linked issue: closes #1050 Enable Native `_SEQUENCE_NUMBER` projections and predicates, including partial Data Evolution updates and primary-key merges. PyPaimon currently falls back to Python for partial-file sequence projections. Java is the reference for the reader contract: `DataEvolutionSplitRead` selects normal-file metadata providers, `ColumnarRowIterator` keeps non-null physical versions and fills nulls from file metadata, and aggregation/partial-update merge functions carry result metadata independently of user value aggregation. ### Brief change log - Register the latest normal-file sequence provider for Data Evolution, including files whose `write_cols` omit the physical column. Preserve that provider through user-column scan pruning, alongside the existing deletion-vector anchor. - Resolve metadata predicates by reserved name; keep them out of positional partition/statistics binding and evaluate exact residuals after tracking metadata is assigned or primary-key merging finishes. - Read predicate-only sequence values through the existing KV metadata column without introducing duplicate columns. - Keep aggregation and partial-update result sequence independent of user aggregates and DELETE/UPDATE_BEFORE retractions. Normalize result row kinds and preserve sequence-group mode for metadata-only projections. - Expose canonical tracking metadata to Python predicate conversion and reject user/metadata name collisions in case-insensitive reads. - Remove `CoreOptions::scan_ignore_lost_file` and the `scan.ignore-lost-files` ROW-sidecar fallback. A missing selected sidecar fails explicitly. The scan pruning change preserves the metadata provider required by Java's reader contract; it does not copy Java's current user-column pruning gap. ### Tests - New persisted integration coverage for metadata-only/mixed projections, repeated partial updates, projected/unprojected predicates, mixed boolean predicates, physical version null filling, BLOB provider isolation, all four PK merge engines, and aggregation/partial-update retractions both within one file and across files. - Strict missing/corrupt ROW-sidecar tests and case-insensitive projection collision coverage. - `cargo test --locked -p paimon --all-targets --features fulltext,vortex`. - `cargo clippy --locked --all-targets --workspace --features fulltext,vortex -- -D warnings` and `cargo fmt --all -- --check`. - Local PyPaimon regression with this binding: 1,413 passed, 31 skipped. Final focused Native tests: 271 passed; normal `py36` regressions: 79 passed on Python 3.13. ### API and Format Existing projection/filter APIs gain `_SEQUENCE_NUMBER` support. The unsupported lost-file option getter is removed. No storage-format or dependency changes. ### Documentation Rust API comments document tracking metadata handling. The paired PyPaimon PR removes the capability fallback and tests Native reads without allowing Python fallback. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
