Zouxxyy opened a new pull request, #10076: URL: https://github.com/apache/paimon/pull/10076
### Purpose Spark 4.1 enables V2 bucketing by default. A plain Paimon filter/project/TopN query can consequently collapse many splits in the same bucket into one scan task. The existing Paimon auto-disable rule only runs during AQE preparation and can invalidate distribution or ordering assumptions when it rewrites an already planned scan. - Add `scan.preserve-data-grouping` (default `false`), supporting table/read options and the `spark.paimon.scan.preserve-data-grouping` session configuration. - Fix the effective layout when creating the scan; retain it through copies and equality. Remove `DisableUnnecessaryPaimonBucketedScan` and its AQE registration. - Use regular split packing by default. Explicit grouped mode requires Spark V2 bucketing and a supported table layout, preserves whole DataSplits, and provides multiple input partitions per bucket for partially clustered joins. **Compatibility:** Workloads relying on bucket distribution to avoid shuffles must now explicitly enable the Paimon option. Grouped scans no longer automatically exit grouped mode when it is unhelpful, and ordinary queries may still have reduced parallelism after opting in. The former `spark.sql.sources.bucketing.autoBucketedScan.enabled` setting no longer changes Paimon scan layout. Configuration and migration documentation describe these changes. ### Tests - 60 regression executions passed: Spark 4.1.2 (39) and 4.0.3 (21), covering scan parallelism, configuration precedence, scan copies/equality, sorting, aggregation, SPJ/partial clustering, updates/deletes and whole-split packing. - 32 additional local review executions passed: Spark 3.3.4 (5), 3.5.8 (21), and 4.1.2 (6). Temporary review fixtures are not included as permanent suites. - Full package builds passed with checks enabled; Spark 3.3/3.5 used JDK 8 and a clean build. Produced JARs were checked for the current scan classes and removal of the old rule. - Spark 3.2/3.4 runtime validation and a rerun of the original customer workload with this patch remain outstanding. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
