jackylee-ch commented on PR #934:
URL: https://github.com/apache/paimon-rust/pull/934#issuecomment-5912447141
Addressed, and rebased onto current main — the branch now sits on top of the
merged `$aggregation_fields` table, and both system tables register side by
side.
`collect_bucket_rows` no longer groups by the rendered partition string.
`aggregate_bucket_rows` now keys on the serialized `BinaryRow` bytes (which are
injective), keeping the rendered string only for output and ordering. So
`(p1='a, b', p2='c')` and `(p1='a', p2='b, c')` stay separate even though both
render `{a, b, c}`. This matches Java, which groups by the binary partition
before formatting (as you noted for `BucketEntry`); the rows still sort by
partition string then bucket, with the raw bytes as a deterministic tie-break
for look-alikes.
Regression: `look_alike_partitions_do_not_merge` builds exactly those two
partitions, asserts the precondition that their rendered strings collide (so
string keying would merge them), and asserts `aggregate_bucket_rows` returns 2
rows. I verified it is non-vacuous: keying on the string alone makes it fail
with `rows.len()` == 1; restoring the byte key passes. I placed the regression
at the aggregation boundary rather than SQL so the collision is exercised
deterministically, without depending on how a comma-bearing partition value
round-trips through INSERT.
The `$aggregation_fields` merge only touched the shared registration list
and the test-file tail, so the rebase was mechanical. `system_tables` (24
tests, including `$aggregation_fields` and `$buckets`) and the `buckets` unit
tests pass; `clippy -p paimon-datafusion --all-targets --features
fulltext,vortex -D warnings` is clean.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]