[
https://issues.apache.org/jira/browse/SPARK-59993?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Haotian Sun updated SPARK-59993:
--------------------------------
Description:
Collect Python worker metrics with a collector shared within each worker
process and reset at task entry. Decode numeric worker reports through the
generic Scala reader while preserving lifecycle timing and spill accounting.
For modern Arrow-optimized scalar Python UDFs (SQL_ARROW_BATCHED_UDF with
non-legacy conversion), record data read, input preparation, UDF execution,
output preparation, and data write durations, plus timed task and batch counts.
Record the count of pipelined Python worker tasks so analyses can identify
operators whose read and UDF intervals may overlap. Register the new values as
SQL metrics for the Spark UI and QPL pipeline. These are internal metrics with
no user-facing behavior change.
was:
Add task-local timing accumulation in the Python worker and report input
conversion, UDF execution, and output conversion durations for scalar pandas
UDFs, together with task and batch coverage counters.
Decode numeric worker reports into a generic map and register the new SQL
metrics so their values are available in the Spark SQL UI. Keep lifecycle
timing calculations and metric updates in the Python runner.
Summary: Collect worker and Arrow-optimized Python UDF timing metrics
(was: Collect scalar pandas UDF phase timing metrics)
> Collect worker and Arrow-optimized Python UDF timing metrics
> ------------------------------------------------------------
>
> Key: SPARK-59993
> URL: https://issues.apache.org/jira/browse/SPARK-59993
> Project: Spark
> Issue Type: Sub-task
> Components: PySpark
> Affects Versions: 4.4.0
> Reporter: Haotian Sun
> Priority: Major
> Labels: pull-request-available
>
> Collect Python worker metrics with a collector shared within each worker
> process and reset at task entry. Decode numeric worker reports through the
> generic Scala reader while preserving lifecycle timing and spill accounting.
> For modern Arrow-optimized scalar Python UDFs (SQL_ARROW_BATCHED_UDF with
> non-legacy conversion), record data read, input preparation, UDF execution,
> output preparation, and data write durations, plus timed task and batch
> counts. Record the count of pipelined Python worker tasks so analyses can
> identify operators whose read and UDF intervals may overlap. Register the new
> values as SQL metrics for the Spark UI and QPL pipeline. These are internal
> metrics with no user-facing behavior change.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]