jackylee-ch commented on PR #934:
URL: https://github.com/apache/paimon-rust/pull/934#issuecomment-5912447141

   Addressed, and rebased onto current main — the branch now sits on top of the 
merged `$aggregation_fields` table, and both system tables register side by 
side.
   
   `collect_bucket_rows` no longer groups by the rendered partition string. 
`aggregate_bucket_rows` now keys on the serialized `BinaryRow` bytes (which are 
injective), keeping the rendered string only for output and ordering. So 
`(p1='a, b', p2='c')` and `(p1='a', p2='b, c')` stay separate even though both 
render `{a, b, c}`. This matches Java, which groups by the binary partition 
before formatting (as you noted for `BucketEntry`); the rows still sort by 
partition string then bucket, with the raw bytes as a deterministic tie-break 
for look-alikes.
   
   Regression: `look_alike_partitions_do_not_merge` builds exactly those two 
partitions, asserts the precondition that their rendered strings collide (so 
string keying would merge them), and asserts `aggregate_bucket_rows` returns 2 
rows. I verified it is non-vacuous: keying on the string alone makes it fail 
with `rows.len()` == 1; restoring the byte key passes. I placed the regression 
at the aggregation boundary rather than SQL so the collision is exercised 
deterministically, without depending on how a comma-bearing partition value 
round-trips through INSERT.
   
   The `$aggregation_fields` merge only touched the shared registration list 
and the test-file tail, so the rebase was mechanical. `system_tables` (24 
tests, including `$aggregation_fields` and `$buckets`) and the `buckets` unit 
tests pass; `clippy -p paimon-datafusion --all-targets --features 
fulltext,vortex -D warnings` is clean.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to