JingsongLi opened a new pull request, #1061:
URL: https://github.com/apache/paimon-rust/pull/1061

   ### Purpose
   
   Closes #1060.
   
   Enable PyPaimon's existing named-row upsert capability to use Rust when 
source rows supply different non-key fields. Rust currently rejects batches 
with different schemas, forcing the Python adapter to fall back.
   
   An absent field must remain absent rather than becoming an explicit NULL. 
Matched rows write the requested update projection; unmatched rows append their 
actual supplied fields. Superseded source rows must not constrain the winning 
row's update or append columns.
   
   ### Brief change log
   
   - Normalize each source batch independently and concatenate only matching 
keys.
   - Preserve source positions and row-ID alignment across batch boundaries and 
duplicate target rows.
   - Validate matched update columns after last-write-wins deduplication.
   - Validate surviving append field sets before staging files. Appends in one 
partition share a write type; different partitions may use different write 
types.
   - Allow the existing Python binding upsert methods to accept either an Arrow 
Table or a sequence of RecordBatch values. Matching, deduplication and column 
validation remain in Rust core.
   
   Java has no equivalent named-row key-upsert API. The writes reuse its 
underlying projection semantics: 
[TableWriteImpl.withWriteType](https://github.com/apache/paimon/blob/08122eef82/paimon-core/src/main/java/org/apache/paimon/table/sink/TableWriteImpl.java#L135)
 and 
[BaseAppendFileStoreWrite.withWriteType](https://github.com/apache/paimon/blob/08122eef82/paimon-core/src/main/java/org/apache/paimon/operation/BaseAppendFileStoreWrite.java#L208).
   
   ### Tests
   
   - Baseline reproduction failed with `Arrow batches in one upsert input must 
have the same columns` before the fix.
   - Core library: 3745 passed, 6 ignored.
   - Update integration suites: 46 passed (`table_update_test`, 
`table_update_nested_test`, `table_update_paths_test`).
   - Python binding suite: 321 passed.
   - Companion PyPaimon update/upsert regression suites using this local 
binding: 991 passed, 2 skipped, 36 subtests passed. Native execution counters: 
1112 plans, 1345 reads, 760 writes.
   - Clippy 1.98.0 passed for `paimon` and `pypaimon_rust`, all targets with 
warnings denied; rustfmt and diff checks passed.
   
   New cases cover differing field sets, last-write-wins across batches, 
fan-out to duplicate targets, explicit NULL, partial append metadata, implicit 
partition keys, differing projections across partitions and validation before 
staging.
   
   ### API and Format
   
   The Rust core API signature is unchanged. Existing Python binding upsert 
methods additionally accept `Sequence[pyarrow.RecordBatch]`. No storage format 
or dependency changes.
   
   ### Documentation
   
   Updated core API documentation and Python binding type stubs to describe 
field presence and per-partition append validation.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to