FelixYBW commented on issue #13140: URL: https://github.com/apache/gluten/issues/13140#issuecomment-5844213039
## What Gluten can reuse from Spark, by use case ### 1. Wrap: C Data import → `loadColumns` - **Gluten today** ([ArrowWritableColumnVector](https://github.com/apache/gluten/blob/main/gluten-cbo/src/main/scala/org/apache/gluten/planner/cost/GlutenCostModel.scala)): `ArrowAbiUtil`, `ColumnarBatches.load` - **Reuse in Spark**: Import into a `VectorSchemaRoot`, then `new ArrowColumnVector(v)` per vector. `PythonArrowOutput` and `ArrowCachedBatchSerializer` do exactly this. - **Origin**: `ArrowColumnVector` (Spark 2.3, SPARK-21472); vector access added in SPARK-38028 (3.3) - **Callable from Gluten?**: ✅ Yes, public --- ### 2. Write rows: `put*` in `ColumnarRangeExec`, `ColumnarPartialProjectExec` / Generate, `ArrowColumnarRow`, `ExecUtil` - **Reuse in Spark**: `ArrowWriter.create(root)`, `write(row)`, `finish()`, then wrap the root's vectors. The Arrow cache uses the same pattern. - **Origin**: `ArrowWriter` (Spark 2.3, SPARK-21440) - **Callable from Gluten?**: ✅ Yes, public --- ### 3. Unwrap for export: `ArrowWritableColumnVector` → C Data export (`ColumnarBatches.offload`) - **Reuse in Spark**: `getValueVector`, `VectorSchemaRoot.of(...)`, then export via Arrow's `Data.exportVectorSchemaRoot` - **Origin**: [#55120](https://github.com/apache/spark/pull/55120)'s `writeArrowDirect` pattern - **Callable from Gluten?**: ⚠️ Pattern only — the method is private --- ### 4. Arrow → IPC for Python: the Gluten runner's `VectorLoader` into its own root - **Reuse in Spark**: `VectorSchemaRoot.of` + `VectorUnloader` + `MessageSerializer.serialize` - **Origin**: [#55120](https://github.com/apache/spark/pull/55120)'s `writeArrowDirect` - **Callable from Gluten?**: ⚠️ Pattern only; on Spark 4.2+ Gluten doesn't need its own Python exec at all --- ### 5. Layout safety: none today (Gluten assumes Velox exports Spark's layout) - **Reuse in Spark**: `ArrowUtils.isCompatibleWithDeclaredField(actual, declared)` - **Origin**: Used by [#55120](https://github.com/apache/spark/pull/55120)'s `isArrowBacked` (added in master after #55120) - **Callable from Gluten?**: ✅ `private[sql]` — reachable from Gluten's `org.apache.spark.sql` packages --- ### 6. Pass-through columns: keeping `ColumnVector` references to join with results - **Reuse in Spark**: Same approach, in `ColumnarArrowEvalPythonEvaluatorFactory.evalArrowColumnar` - **Origin**: [#55120](https://github.com/apache/spark/pull/55120) - **Callable from Gluten?**: ⚠️ Pattern only (private) --- ## What doesn't carry over - **Reference counting and in-place offload/load.** Nothing in Spark does this. `ArrowColumnVector.close()` just closes the vector, so Gluten has to move to a "producer owns the batch; to keep it, transfer the vectors" model. The prototype's transitions show that pattern. - **Gluten's allocator.** `ArrowWriter.create(root)` accepts a root allocated from Gluten's managed allocator, so memory stays tracked. Vectors that arrive from Spark's own allocators (the Python output, the Arrow cache) still need the copy-or-transfer rule described above. --- ## A small, worthwhile upstream refactor Move [#55120](https://github.com/apache/spark/pull/55120)'s two private helpers into a shared public utility, so Gluten, the Arrow cache, and the UDTF path stop duplicating them: 1. **"Is this batch Arrow with the declared layout"** — `isArrowBacked` + `isCompatibleWithDeclaredField` 2. **"Unwrap an `ArrowColumnVector` batch into a `VectorSchemaRoot` or record batch"** — `writeArrowDirect` This fits naturally next to SPARK-37124's row-to-Arrow operator, which would reuse `ArrowWriter` the same way. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
