viirya commented on PR #58978: URL: https://github.com/apache/spark/pull/58978#issuecomment-5804693071
Thanks for tracing these cases. I've addressed them in 13ccfb555c0 and added regression coverage. I took your suggestion to use the existing `PythonUDF` + `ArrowEvalPython` path. The public `@inprocess_udf` API remains, but it now creates a regular `PythonUDF` with a new `SQL_SCALAR_ARROW_INPROCESS_UDF` eval type. Physical planning selects `InProcessArrowEvalExec`. This removes the separate expression, logical node, and extraction rule, and brings aggregation, nondeterminism, semantic equality, Python UDF guards, and filter/limit pushdown under the existing contracts. For the runtime issues, sliced results are normalized before CDI export, and both deserialization and invocation convert `BaseException`, including `SystemExit`, into task failures. Only UDF arguments go through Arrow; original rows stay in `HybridRowQueue`. Each batch has fresh input buffers so later batches cannot overwrite arrays retained by Python. Functions are registered once per task and released at task completion. I also addressed the smaller items: added `producedAttributes`; reused the existing non-excludable extraction rule; made plugin initialization log `LinkageError` as well as `NonFatal` failures; corrected the Python/PyArrow requirements, Docker example, and archive layout; and removed the em-dashes from the source/POM. JEP is now 4.3.2. Local validation passed: the Hive-enabled package build, SQL test compilation, 27 Scala tests across the in-process and extraction suites, 19 Python runtime tests, and the Python integration suite with JEP 4.3.2. The integration suite runs with `local[2]` and `spark.task.cpus=0.5`. Scala style checks and Python Ruff checks passed too. I also checked ordinary SQL, Python UDFs, and worker Arrow UDFs without the optional JEP/CDI jars. I have not rerun performance benchmarks for this revision. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
