akshaytayal opened a new issue, #13139:
URL: https://github.com/apache/gluten/issues/13139

   **Umbrella:** GLUTEN-12569 (Spark 4.2 support)
   
   ### Background
   Spark 4.2 restructured Storage-Partitioned Join (SPJ) partitioning. Verified 
against the Spark jars:
   - `GroupPartitionsExec` is **new in Spark 4.2.0** (absent in 4.1.1). In 4.2 
the scan (`BatchScanExec`) emits **ungrouped** partitions — one per split, 
duplicate keys retained — and grouping is deferred to `GroupPartitionsExec`. In 
4.1 the scan produced the key-grouped layout itself.
   
   ### Problem
   The Spark 4.2 shim 
(`shims/spark42/src/main/scala/org/apache/spark/sql/execution/datasources/v2/BatchScanExecShim.scala`,
 `filteredPartitions`) still rebuilds the old 4.1 layout via 
`groupBy(InternalRowComparableWrapper(key))`, collapsing duplicate-key splits. 
For keys `[A, A, B]` it advertises 3 partitions but builds only 2 native RDD 
partitions; Spark 4.2's `GroupPartitionsExec` then derives parent indices 
assuming the ungrouped count → index mismatch / wrong per-key multiplicity.
   
   ### Scope / why deferred
   Only reachable via scans reporting `KeyGroupedPartitioning` (DSv2 
`SupportsReportPartitioning`). Verified `FileScan` (Parquet/ORC/CSV) does 
**not** implement it, and Delta does not use DSv2 SPJ. The only native consumer 
is the **Iceberg transformer**, which is **not enabled** on Spark 4.2 
(dependency unavailable). It is therefore currently **dormant / 
non-triggerable** and cannot be behaviorally tested yet.
   
   ### Required work (when enabling native keyed DSv2 / Iceberg on Spark 4.2)
   1. Preserve one slot per original sorted split (pad filtered-out splits) and 
let `GroupPartitionsExec` perform grouping — target layout:
      ```
      original [A1,A2,B1]        -> [[A1],[A2],[B1]]
      filter removes A2         -> [[A1],[],[B1]]
      filter removes A1 and A2  -> [[],[],[B1]]
      ```
      *or* explicitly fall native keyed scans back to vanilla Spark until (1) 
is implemented.
   2. Add SPJ tests over an **actually offloaded** scan with partition-count 
and result assertions (including a nondefault timezone and worker reuse) — 
belongs with the `gluten-ut/spark42` UT module.
   
   ### Reference
   Deferred from review discussion on #13126.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to