JingsongLi opened a new pull request, #10368: URL: https://github.com/apache/paimon/pull/10368
### Purpose Align Python pending-bucket event writes with Java and enable the native postpone merge engines once https://github.com/apache/paimon-rust/pull/1020 is merged. Follow-up to #10354. **Dependency:** keep this PR draft until paimon-rust #1020 reaches main. The native CI continues installing `apache/paimon-rust@main`; it does not pin a fork or feature branch. Local native results below use the rebuilt companion Rust change. - Python postpone writes preserve arrival order and duplicate keys rather than sorting/folding them. They validate the first retract, retain its state across checkpoints, populate retract counts/sequence metadata, and hand off prepared files only after all partitions prepare successfully. - Honor `rowkind.field` and apply `ignore-delete` / `ignore-update-before` before bucket routing for both Arrow and row inputs. - Use the existing field aggregators in the Python aggregation writer instead of silently deduplicating the write buffer. For example, two same-key values 10 and 30 now produce 40 with `sum`. - Enable Rust postpone writes for the other merge engines and managed scalar/ARRAY/MAP Blob fields, including fixed-bucket postpone writes. Managed Blob storage and resolution remain entirely in Rust. The tests create schemas with Rust, write through NativeTableWrite and fail if reads fall back to Python. - Keep append Blob deferral separate from Rust PK managed packs, so a payload projection plus LIMIT can stay native. Reject a Python PK Blob fallback before it can emit an append-file layout. No Python managed Blob implementation or schema-validation implementation is added. No CI Rust source change. ### Tests - Eight native writer suites with native plan/read/write/commit enabled: **257 passed** (150 native plans, 226 reads, 450 writes and 50 commits observed). - Python Blob/schema/aggregation/nested-schema/write regression suites: **574 passed**, 1 skipped, 112 subtests passed. - Final focused postpone checkpoint ownership and writer regression run: **104 passed**, 6 subtests passed. - `python -m flake8 --config dev/cfg.ini pypaimon`: passed. Regression cases cover unsorted duplicates, row-kind filters, unsupported retract cleanup, repeated checkpoints, failure after another partition prepared, managed Blob nulls/empty values/duplicate map keys, copying referenced payloads, and payload filtering with projection and LIMIT. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
