JingsongLi opened a new pull request, #10368:
URL: https://github.com/apache/paimon/pull/10368

   ### Purpose
   
   Align Python pending-bucket event writes with Java and enable the native 
postpone merge engines once https://github.com/apache/paimon-rust/pull/1020 is 
merged. Follow-up to #10354.
   
   **Dependency:** keep this PR draft until paimon-rust #1020 reaches main. The 
native CI continues installing `apache/paimon-rust@main`; it does not pin a 
fork or feature branch. Local native results below use the rebuilt companion 
Rust change.
   
   - Python postpone writes preserve arrival order and duplicate keys rather 
than sorting/folding them. They validate the first retract, retain its state 
across checkpoints, populate retract counts/sequence metadata, and hand off 
prepared files only after all partitions prepare successfully.
   - Honor `rowkind.field` and apply `ignore-delete` / `ignore-update-before` 
before bucket routing for both Arrow and row inputs.
   - Use the existing field aggregators in the Python aggregation writer 
instead of silently deduplicating the write buffer. For example, two same-key 
values 10 and 30 now produce 40 with `sum`.
   - Enable Rust postpone writes for the other merge engines and managed 
scalar/ARRAY/MAP Blob fields, including fixed-bucket postpone writes. Managed 
Blob storage and resolution remain entirely in Rust. The tests create schemas 
with Rust, write through NativeTableWrite and fail if reads fall back to Python.
   - Keep append Blob deferral separate from Rust PK managed packs, so a 
payload projection plus LIMIT can stay native. Reject a Python PK Blob fallback 
before it can emit an append-file layout.
   
   No Python managed Blob implementation or schema-validation implementation is 
added. No CI Rust source change.
   
   ### Tests
   
   - Eight native writer suites with native plan/read/write/commit enabled: 
**257 passed** (150 native plans, 226 reads, 450 writes and 50 commits 
observed).
   - Python Blob/schema/aggregation/nested-schema/write regression suites: 
**574 passed**, 1 skipped, 112 subtests passed.
   - Final focused postpone checkpoint ownership and writer regression run: 
**104 passed**, 6 subtests passed.
   - `python -m flake8 --config dev/cfg.ini pypaimon`: passed.
   
   Regression cases cover unsorted duplicates, row-kind filters, unsupported 
retract cleanup, repeated checkpoints, failure after another partition 
prepared, managed Blob nulls/empty values/duplicate map keys, copying 
referenced payloads, and payload filtering with projection and LIMIT.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to