XiDuo You created SPARK-58988:
---------------------------------

             Summary: Fix PartitioningCollection invariant violation in 
subset-key storage-partitioned join
                 Key: SPARK-58988
                 URL: https://issues.apache.org/jira/browse/SPARK-58988
             Project: Spark
          Issue Type: Bug
          Components: SQL
    Affects Versions: 5.0.0
            Reporter: XiDuo You


This is a follow-up to SPARK-56877, which added a check in 
`PartitioningCollection.fromPartitionings` requiring all `KeyedPartitioning`s 
to share equal `partitionKeys`.

When a storage-partitioned join has join keys that are a subset of the 
partition keys (e.g. a v2 table partitioned by `[dt, bucket(16, c1)]` joined on 
`c1` only), and `spark.sql.sources.v2.bucketing.shuffle.enabled` is on so only 
the non-keyed side is re-shuffled, `KeyedPartitioning.createShuffleSpec` 
projects the keyed side to the join keys but keeps them in first-occurrence 
order (`distinct` without sorting). The keyed side, however, goes through 
`GroupPartitionsExec`, which sorts the projected keys. The two sides then carry 
the same keys in a different order, and 
`PartitioningCollection.fromPartitionings` throws:

```
java.lang.IllegalArgumentException: requirement failed: All KeyedPartitionings 
in a PartitioningCollection must have equal partitionKeys
```

The fix sorts the distinct projected keys in `createShuffleSpec` (using 
`toGrouped`, which sorts by the same natural ascending ordering as 
`GroupPartitionsExec`), aligning the two sides.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to