Bryan Cutler created SPARK-28128:
------------------------------------
Summary: Pandas Grouped UDFs should skip over empty partitions
Key: SPARK-28128
URL: https://issues.apache.org/jira/browse/SPARK-28128
Project: Spark
Issue Type: Improvement
Components: PySpark, SQL
Affects Versions: 2.4.3
Reporter: Bryan Cutler
When running FlatMapGroupsInPandasExec or AggregateInPandasExec the shuffle
uses a default number of partitions of 200 in "spark.sql.shuffle.partitions".
If the data is small, e.g. in testing, many of the partitions will be empty but
are treated just the same. For example, ArrowPythonRunner.compute is called and
starts a number of threads that do nothing since there is no iteration. These
computations could be skipped for empty partitions, which will save time
overall.
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]