FelixYBW commented on issue #13140:
URL: https://github.com/apache/gluten/issues/13140#issuecomment-5844213039

   ## What Gluten can reuse from Spark, by use case
   
   ### 1. Wrap: C Data import → `loadColumns`
   - **Gluten today** 
([ArrowWritableColumnVector](https://github.com/apache/gluten/blob/main/gluten-cbo/src/main/scala/org/apache/gluten/planner/cost/GlutenCostModel.scala)):
 `ArrowAbiUtil`, `ColumnarBatches.load`
   - **Reuse in Spark**: Import into a `VectorSchemaRoot`, then `new 
ArrowColumnVector(v)` per vector. `PythonArrowOutput` and 
`ArrowCachedBatchSerializer` do exactly this.
   - **Origin**: `ArrowColumnVector` (Spark 2.3, SPARK-21472); vector access 
added in SPARK-38028 (3.3)
   - **Callable from Gluten?**: ✅ Yes, public
   
   ---
   
   ### 2. Write rows: `put*` in `ColumnarRangeExec`, 
`ColumnarPartialProjectExec` / Generate, `ArrowColumnarRow`, `ExecUtil`
   - **Reuse in Spark**: `ArrowWriter.create(root)`, `write(row)`, `finish()`, 
then wrap the root's vectors. The Arrow cache uses the same pattern.
   - **Origin**: `ArrowWriter` (Spark 2.3, SPARK-21440)
   - **Callable from Gluten?**: ✅ Yes, public
   
   ---
   
   ### 3. Unwrap for export: `ArrowWritableColumnVector` → C Data export 
(`ColumnarBatches.offload`)
   - **Reuse in Spark**: `getValueVector`, `VectorSchemaRoot.of(...)`, then 
export via Arrow's `Data.exportVectorSchemaRoot`
   - **Origin**: [#55120](https://github.com/apache/spark/pull/55120)'s 
`writeArrowDirect` pattern
   - **Callable from Gluten?**: ⚠️ Pattern only — the method is private
   
   ---
   
   ### 4. Arrow → IPC for Python: the Gluten runner's `VectorLoader` into its 
own root
   - **Reuse in Spark**: `VectorSchemaRoot.of` + `VectorUnloader` + 
`MessageSerializer.serialize`
   - **Origin**: [#55120](https://github.com/apache/spark/pull/55120)'s 
`writeArrowDirect`
   - **Callable from Gluten?**: ⚠️ Pattern only; on Spark 4.2+ Gluten doesn't 
need its own Python exec at all
   
   ---
   
   ### 5. Layout safety: none today (Gluten assumes Velox exports Spark's 
layout)
   - **Reuse in Spark**: `ArrowUtils.isCompatibleWithDeclaredField(actual, 
declared)`
   - **Origin**: Used by [#55120](https://github.com/apache/spark/pull/55120)'s 
`isArrowBacked` (added in master after #55120)
   - **Callable from Gluten?**: ✅ `private[sql]` — reachable from Gluten's 
`org.apache.spark.sql` packages
   
   ---
   
   ### 6. Pass-through columns: keeping `ColumnVector` references to join with 
results
   - **Reuse in Spark**: Same approach, in 
`ColumnarArrowEvalPythonEvaluatorFactory.evalArrowColumnar`
   - **Origin**: [#55120](https://github.com/apache/spark/pull/55120)
   - **Callable from Gluten?**: ⚠️ Pattern only (private)
   
   ---
   
   ## What doesn't carry over
   
   - **Reference counting and in-place offload/load.** Nothing in Spark does 
this. `ArrowColumnVector.close()` just closes the vector, so Gluten has to move 
to a "producer owns the batch; to keep it, transfer the vectors" model. The 
prototype's transitions show that pattern.
   - **Gluten's allocator.** `ArrowWriter.create(root)` accepts a root 
allocated from Gluten's managed allocator, so memory stays tracked. Vectors 
that arrive from Spark's own allocators (the Python output, the Arrow cache) 
still need the copy-or-transfer rule described above.
   
   ---
   
   ## A small, worthwhile upstream refactor
   
   Move [#55120](https://github.com/apache/spark/pull/55120)'s two private 
helpers into a shared public utility, so Gluten, the Arrow cache, and the UDTF 
path stop duplicating them:
   
   1. **"Is this batch Arrow with the declared layout"** — `isArrowBacked` + 
`isCompatibleWithDeclaredField`
   2. **"Unwrap an `ArrowColumnVector` batch into a `VectorSchemaRoot` or 
record batch"** — `writeArrowDirect`
   
   This fits naturally next to SPARK-37124's row-to-Arrow operator, which would 
reuse `ArrowWriter` the same way.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to