szehon-ho opened a new pull request, #46325:
URL: https://github.com/apache/spark/pull/46325
### What changes were proposed in this pull request?
If spark.sql.v2.bucketing.allowJoinKeysSubsetOfPartitionKeys.enabled is
true, change KeyGroupedPartitioning.satisfies0(distribution) check from all
clustering keys (here, join keys) being in partition keys, to the two sets
overlapping.
### Why are the changes needed?
If spark.sql.v2.bucketing.allowJoinKeysSubsetOfPartitionKeys.enabled is
true, then SPJ no longer triggers if there are more join keys than partition
keys. But SPJ is supported in this case if flag is false.
### Does this PR introduce _any_ user-facing change?
No
### How was this patch tested?
Added tests in KeyGroupedPartitioningSuite
### Was this patch authored or co-authored using generative AI tooling?
No
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]