L. C. Hsieh created SPARK-60071:
-----------------------------------
Summary: Use identity equality for HybridRowQueue
Key: SPARK-60071
URL: https://issues.apache.org/jira/browse/SPARK-60071
Project: Spark
Issue Type: Bug
Components: SQL
Affects Versions: 5.0.0
Reporter: L. C. Hsieh
HybridRowQueue is a case class, so two queues created in the same task with the
same TaskMemoryManager, temp dir, number of fields, SerializerManager and
lockFree flag are equal. Because HybridRowQueue is a MemoryConsumer (via
HybridQueue), this breaks memory management:
1. TaskMemoryManager tracks consumers in a HashSet. A second queue that equals
an already registered one is never added, so it is never offered for spilling
when another consumer needs memory, and its memory is not attributed in the
memory usage breakdown on OOM.
2. HybridQueue.spill(size, trigger) returns 0 when trigger == this, which uses
equals. A queue therefore refuses to spill when an equal but distinct queue
triggers the spill.
This can happen in a real query, e.g. when several partitions evaluated by a
Python UDF node are processed in one task (coalesce), or when both sides of a
sort-merge join contain Python UDF eval nodes with the same input width.
The fix is to give HybridRowQueue identity equality by making it a regular
class.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]