viirya commented on PR #58978:
URL: https://github.com/apache/spark/pull/58978#issuecomment-5804693071

   Thanks for tracing these cases. I've addressed them in 13ccfb555c0 and added 
regression coverage.
   
   I took your suggestion to use the existing `PythonUDF` + `ArrowEvalPython` 
path. The public `@inprocess_udf` API remains, but it now creates a regular 
`PythonUDF` with a new `SQL_SCALAR_ARROW_INPROCESS_UDF` eval type. Physical 
planning selects `InProcessArrowEvalExec`. This removes the separate 
expression, logical node, and extraction rule, and brings aggregation, 
nondeterminism, semantic equality, Python UDF guards, and filter/limit pushdown 
under the existing contracts.
   
   For the runtime issues, sliced results are normalized before CDI export, and 
both deserialization and invocation convert `BaseException`, including 
`SystemExit`, into task failures. Only UDF arguments go through Arrow; original 
rows stay in `HybridRowQueue`. Each batch has fresh input buffers so later 
batches cannot overwrite arrays retained by Python. Functions are registered 
once per task and released at task completion.
   
   I also addressed the smaller items: added `producedAttributes`; reused the 
existing non-excludable extraction rule; made plugin initialization log 
`LinkageError` as well as `NonFatal` failures; corrected the Python/PyArrow 
requirements, Docker example, and archive layout; and removed the em-dashes 
from the source/POM. JEP is now 4.3.2.
   
   Local validation passed: the Hive-enabled package build, SQL test 
compilation, 27 Scala tests across the in-process and extraction suites, 19 
Python runtime tests, and the Python integration suite with JEP 4.3.2. The 
integration suite runs with `local[2]` and `spark.task.cpus=0.5`. Scala style 
checks and Python Ruff checks passed too. I also checked ordinary SQL, Python 
UDFs, and worker Arrow UDFs without the optional JEP/CDI jars. I have not rerun 
performance benchmarks for this revision.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to