minni31 opened a new pull request, #13171: URL: https://github.com/apache/gluten/pull/13171
## What changes are proposed in this pull request? This PR enables native key-grouped columnar shuffle for the Velox backend when Spark's partition metadata can be mapped safely. - Add Spark-version shims for key-grouped shuffle metadata across Spark 3.4 through 4.2. - Compute partition IDs on the JVM and prepend a temporary PID column for the existing native range-shuffle path. - Preserve vanilla Spark shuffle for ClickHouse, Bolt, and duplicate or partially clustered partition metadata. - Keep unknown-key routing compatible with the active Spark runtime, including content-based binary-key hashing. - Add focused Spark 4.0/4.1 integration coverage and a deterministic binary fallback-hash regression. This supersedes #12084 and addresses its review concern that partially clustered or duplicated partition values could make the advertised partition count diverge from the routing map. ## How was this patch tested? - `./dev/format-scala-code.sh apply` - `git diff --check` - License headers validated for all changed files with `.github/workflows/util/license-header.py` - Added Spark 4.0/4.1 key-grouped partitioning regressions and `ExecUtilSuite` The project build and test suites were not run locally; this draft is pending CI and external validation. ## Was this patch authored or co-authored using generative AI tooling? Generated-by: GitHub Copilot CLI 1.0.86 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
