Zouxxyy opened a new pull request, #10076:
URL: https://github.com/apache/paimon/pull/10076

   ### Purpose
   
   Spark 4.1 enables V2 bucketing by default. A plain Paimon 
filter/project/TopN query can consequently collapse many splits in the same 
bucket into one scan task. The existing Paimon auto-disable rule only runs 
during AQE preparation and can invalidate distribution or ordering assumptions 
when it rewrites an already planned scan.
   
   - Add `scan.preserve-data-grouping` (default `false`), supporting table/read 
options and the `spark.paimon.scan.preserve-data-grouping` session 
configuration.
   - Fix the effective layout when creating the scan; retain it through copies 
and equality. Remove `DisableUnnecessaryPaimonBucketedScan` and its AQE 
registration.
   - Use regular split packing by default. Explicit grouped mode requires Spark 
V2 bucketing and a supported table layout, preserves whole DataSplits, and 
provides multiple input partitions per bucket for partially clustered joins.
   
   **Compatibility:** Workloads relying on bucket distribution to avoid 
shuffles must now explicitly enable the Paimon option. Grouped scans no longer 
automatically exit grouped mode when it is unhelpful, and ordinary queries may 
still have reduced parallelism after opting in. The former 
`spark.sql.sources.bucketing.autoBucketedScan.enabled` setting no longer 
changes Paimon scan layout. Configuration and migration documentation describe 
these changes.
   
   ### Tests
   
   - 60 regression executions passed: Spark 4.1.2 (39) and 4.0.3 (21), covering 
scan parallelism, configuration precedence, scan copies/equality, sorting, 
aggregation, SPJ/partial clustering, updates/deletes and whole-split packing.
   - 32 additional local review executions passed: Spark 3.3.4 (5), 3.5.8 (21), 
and 4.1.2 (6). Temporary review fixtures are not included as permanent suites.
   - Full package builds passed with checks enabled; Spark 3.3/3.5 used JDK 8 
and a clean build. Produced JARs were checked for the current scan classes and 
removal of the old rule.
   - Spark 3.2/3.4 runtime validation and a rerun of the original customer 
workload with this patch remain outstanding.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to