jackylee-ch opened a new pull request, #10162:
URL: https://github.com/apache/paimon/pull/10162

   ### Purpose
   
   The read path aggregates aggregation-engine tables (`SplitRead` builds 
`AggregateMergeFunction`), but the write buffer degraded every aggregation 
table to `DeduplicateMergeFunction`. So duplicate keys inside a single 
`write_arrow` were silently reduced to the last row instead of aggregated — a 
batch of `total=10/20/30` for one key produced `30`, not `60`. It only looked 
correct when each row was committed separately, because read re-aggregates 
across files.
   
   `FileStoreWrite._build_pk_merge_function` now builds the same 
`AggregateMergeFunction` as the read path for supported aggregation tables, so 
same-key rows in one buffer are partially aggregated (read and compaction 
re-aggregate across files, matching Java's per-buffer partial aggregation). 
Aggregation configured with options pypaimon does not implement (retract 
opt-ins, sequence groups, out-of-scope aggregators such as `rbm64`) still falls 
back to deduplicate with a warning — the read-side `check_supported` guard 
raises for those, so the user gets the explicit error at read.
   
   ### Tests
   
   `test_aggregation_e2e` aggregates duplicate keys within one `write_arrow` 
(sum → 60, max, default last-non-null; disjoint key untouched) — returns `30` 
on the pre-fix dedupe path. The write-side fallback-warning test is repurposed 
to an unsupported-option table, since supported aggregation no longer falls 
back.
   
   Written with Claude Code; verification is mine.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to