dongjoon-hyun commented on code in PR #58978:
URL: https://github.com/apache/spark/pull/58978#discussion_r4117401026


##########
python/pyspark/inprocess/runtime.py:
##########
@@ -0,0 +1,310 @@
+#
+# Licensed to the Apache Software Foundation (ASF) under one or more
+# contributor license agreements.  See the NOTICE file distributed with
+# this work for additional information regarding copyright ownership.
+# The ASF licenses this file to You under the Apache License, Version 2.0
+# (the "License"); you may not use this file except in compliance with
+# the License.  You may obtain a copy of the License at
+#
+#    http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing, software
+# distributed under the License is distributed on an "AS IS" BASIS,
+# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+# See the License for the specific language governing permissions and
+# limitations under the License.
+#
+
+
+"""Arrow CDI entry points called on the executor's dedicated JEP interpreter 
thread.
+
+Functions are registered once per task and released when that task finishes. 
Calls
+pass only a handle and CDI addresses, so large closures are not copied per 
batch.
+"""
+
+import sys
+from typing import Any, Callable, Iterable, Optional, Sequence
+
+import pyarrow as pa
+import pyarrow.compute as pc
+
+from pyspark import cloudpickle
+from pyspark.errors import PySparkRuntimeError
+from pyspark.sql.pandas.utils import require_minimum_pyarrow_version
+from pyspark.util import _format_exception
+
+_UDF_TRACEBACK_SENTINEL = "__INPROCESS_UDF_TRACEBACK__:"
+NullChecker = Callable[[pa.Array], None]
+_udfs: dict[str, tuple[Callable[..., pa.Array], pa.DataType, NullChecker, 
bool, bool, bool]] = {}
+
+
+def _jep_safe_message(message: str) -> str:
+    # JEP uses JNI modified UTF-8 for exception text. Keep the transport ASCII 
and
+    # escape NUL explicitly; ordinary UTF-8 and embedded NUL are not safe here.
+    return message.encode("ascii", 
"backslashreplace").decode("ascii").replace("\0", "\\x00")

Review Comment:
   **[Low] This escapes more than JNI needs, so non-English error messages 
become unreadable.**
   
   Following up on my earlier comment: JEP builds the message with 
`PyUnicode_AsUTF8` + `NewStringUTF`, and modified UTF-8 is identical to 
standard UTF-8 for every BMP code point except NUL, while lone surrogates can't 
be encoded at all. So only NUL, surrogates and supplementary characters 
(U+10000 and above) actually need escaping. Escaping everything turns `raise 
ValueError("잘못된 값: café")` into `ValueError: \uc798\ubabb\ub41c \uac12: 
caf\xe9` on the driver, escapes non-ASCII source lines and file paths in 
tracebacks (which also shifts the `^^^^` markers), and since backslashes 
themselves aren't escaped, the original text can't be recovered. Worker UDFs 
show the text as is.
   
   Suggestion: escape only the unsafe subset, e.g. 
`re.sub("[\x00\ud800-\udfff\U00010000-\U0010ffff]", lambda m: 
m.group().encode("unicode_escape").decode("ascii"), message)`. 
`test_exception_text_is_safe_for_jni` and 
`test_exception_unicode_and_nul_survive_jep` still pass with that; the guide 
sentence and this comment would need a small update.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to