dongjoon-hyun commented on PR #58978:
URL: https://github.com/apache/spark/pull/58978#issuecomment-5819215703
Thank you for working on this, @viirya. In-process execution over the Arrow
C Data Interface is an interesting direction for cutting the JVM/Python
transfer cost, and the PR is careful about CDI ownership, cancellation, and
resource cleanup.
I left 15 inline comments. Here is a summary:
**Result type validation rejects valid results**
- The exact Arrow `Field` comparison in `cdiToColumn` fails for struct
fields with Spark metadata, because the JVM writes compact JSON and Python
writes JSON with spaces.
- It also fails for `TimeType` and the nanos timestamp types, because the
precision metadata cannot be carried by `pa.Array._export_to_c`.
- `Variant`/`Geometry`/`Geography` results can never pass when
`useLargeVarTypes=true` (`binary` vs `LargeBinary`).
**Behavior differences from Python workers**
- `PYTHONHASHSEED` is not fixed on cluster executors, so deterministic UDFs
that use `hash()` can differ across executors.
- `hideTraceback` and `simplifiedTraceback` are ignored.
- Spark Connect can plan eval type 258 in-process, and `register` on Connect
accepts the wrapper, which then fails later at call time.
**Robustness**
- A UDF that never returns blocks all in-process work on the executor.
- Re-initialization in the same JVM fails with a misleading installation
error while the old session is still stopping.
- A `SystemExit` during interpreter initialization is not guarded.
**Benchmark, docs, and minor items**
- The ASV benchmark never registers the plugin, so the `inprocess` cases
always fail.
- The YARN/K8s examples do not put the JEP and `arrow-c-data` JARs on the
classpath, and `python3 -c 'import jep'` fails in standalone Python.
- Minor items:
- The timing metrics are truncated to milliseconds on every call.
- DDL string return types fail late.
- `_check_nested_nulls` copies nested results on every batch.
- `bridge.py` is an unused module.
Thank you again for the contribution. I'm looking forward to the next
revision.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]