jackylee-ch opened a new pull request, #10162: URL: https://github.com/apache/paimon/pull/10162
### Purpose The read path aggregates aggregation-engine tables (`SplitRead` builds `AggregateMergeFunction`), but the write buffer degraded every aggregation table to `DeduplicateMergeFunction`. So duplicate keys inside a single `write_arrow` were silently reduced to the last row instead of aggregated — a batch of `total=10/20/30` for one key produced `30`, not `60`. It only looked correct when each row was committed separately, because read re-aggregates across files. `FileStoreWrite._build_pk_merge_function` now builds the same `AggregateMergeFunction` as the read path for supported aggregation tables, so same-key rows in one buffer are partially aggregated (read and compaction re-aggregate across files, matching Java's per-buffer partial aggregation). Aggregation configured with options pypaimon does not implement (retract opt-ins, sequence groups, out-of-scope aggregators such as `rbm64`) still falls back to deduplicate with a warning — the read-side `check_supported` guard raises for those, so the user gets the explicit error at read. ### Tests `test_aggregation_e2e` aggregates duplicate keys within one `write_arrow` (sum → 60, max, default last-non-null; disjoint key untouched) — returns `30` on the pre-fix dedupe path. The write-side fallback-warning test is repurposed to an unsupported-option table, since supported aggregation no longer falls back. Written with Claude Code; verification is mine. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
