JingsongLi commented on PR #934:
URL: https://github.com/apache/paimon-rust/pull/934#issuecomment-5831712177
Requirement fit: SUPPORTED; bucket skew diagnostics are useful. However,
**P1: group by the raw partition row, not its rendered string.**
`collect_bucket_rows` uses `(Option<String>, i32)` as the map key. The Java row
formatter is not injective: partitions `(p1='a, b', p2='c')` and `(p1='a',
p2='b, c')` both render `{a, b, c}`. I reproduced this through DataFusion SQL
at `3551282f` with a two-column partitioned table (`bucket=1`,
`bucket-key=id`): inserting one row in each partition yields 2 `$files` rows
but only 1 `$buckets` row. The new table silently combines unrelated partitions
and reports incorrect counts/sizes. Please group by `BinaryRow` plus bucket,
then render the partition only for output; add this two-partition regression.
The original `$buckets` end-to-end test, query-authorization fail-closed
test, formatting, and diff checks pass. I restored the temporary reproducer
after running it. Java `BucketsTable` / `BucketEntry` group by the binary
partition before formatting:
https://github.com/apache/paimon/blob/master/paimon-core/src/main/java/org/apache/paimon/manifest/BucketEntry.java
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]