JingsongLi opened a new issue, #1060:
URL: https://github.com/apache/paimon-rust/issues/1060
### Motivation
PyPaimon's existing `upsert_by_key` API accepts named `GenericRow` inputs
with different non-key field sets. Its native adapter currently falls back to
Python for these inputs because Rust requires every source batch to have the
same schema.
For example, with `with_update_type(["value"])`, an input can contain a
matched row `{id: 1, value: 111}` and an unmatched row `{id: 4, score: 400}`.
The matched row updates `value`; the unmatched row appends only its supplied
fields. Padding absent fields with NULL loses the distinction between field
absence and an explicit NULL.
### Proposed change
- Preserve source batch schemas and match using key-only batches.
- Deduplicate keys before validating surviving matched and appended rows.
- Require surviving appended rows in each partition to share a field set;
allow different partitions to use different write types.
- Extend the existing Python binding upsert method to accept a sequence of
Arrow RecordBatch values.
Reuse Java's append write type and Data Evolution update projection
semantics. Keep matching, winner selection and validation in Rust core. No
storage format change or new PyPaimon public API is needed.
### Willingness to contribute
I have a tested implementation ready.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]