JingsongLi commented on PR #934:
URL: https://github.com/apache/paimon-rust/pull/934#issuecomment-5831712177

   Requirement fit: SUPPORTED; bucket skew diagnostics are useful. However, 
**P1: group by the raw partition row, not its rendered string.** 
`collect_bucket_rows` uses `(Option<String>, i32)` as the map key. The Java row 
formatter is not injective: partitions `(p1='a, b', p2='c')` and `(p1='a', 
p2='b, c')` both render `{a, b, c}`. I reproduced this through DataFusion SQL 
at `3551282f` with a two-column partitioned table (`bucket=1`, 
`bucket-key=id`): inserting one row in each partition yields 2 `$files` rows 
but only 1 `$buckets` row. The new table silently combines unrelated partitions 
and reports incorrect counts/sizes. Please group by `BinaryRow` plus bucket, 
then render the partition only for output; add this two-partition regression.
   
   The original `$buckets` end-to-end test, query-authorization fail-closed 
test, formatting, and diff checks pass. I restored the temporary reproducer 
after running it. Java `BucketsTable` / `BucketEntry` group by the binary 
partition before formatting: 
https://github.com/apache/paimon/blob/master/paimon-core/src/main/java/org/apache/paimon/manifest/BucketEntry.java
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to