viirya opened a new issue, #6812:
URL: https://github.com/apache/datafusion-comet/issues/6812

   #6403 made `CometUnifiedShuffleMemoryAllocator` retry a page allocation that 
Spark fails with `NoSuchElementException: key not found: <taskAttemptId>`. That 
happens when a request waits in `ExecutionMemoryPool` below the task's minimum 
share, and another consumer of the task releases the task's last bytes in the 
meantime ([SPARK-59444](https://issues.apache.org/jira/browse/SPARK-59444)).
   
   Comet's native memory pools reach Spark through 
`CometTaskMemoryManager.acquireMemory`, which calls 
`TaskMemoryManager.acquireExecutionMemory` without that guard. A native 
`try_grow` waiting in Spark gets the exception back as a `CometError` and the 
task fails. The other consumer can be another native plan in the same task or a 
JVM consumer such as the shuffle allocator. `grow` logs it and treats the 
request as overcommit.
   
   This reproduces with a real `UnifiedMemoryManager`. With a 100 byte off-heap 
pool, another task holds 90 bytes and the task holds 10 bytes through a second 
consumer. A native `acquireMemory(20)` waits below the minimum share of 25. 
When the second consumer releases its 10 bytes, the waiting request wakes to 
`key not found: 0`.
   
   The Spark fix, 
[apache/spark#58747](https://github.com/apache/spark/pull/58747), is still open 
against master. Even once it merges, the Spark versions Comet supports will not 
have it, so Comet needs its own guard as it did in #6403.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to