viirya opened a new issue, #6812: URL: https://github.com/apache/datafusion-comet/issues/6812
#6403 made `CometUnifiedShuffleMemoryAllocator` retry a page allocation that Spark fails with `NoSuchElementException: key not found: <taskAttemptId>`. That happens when a request waits in `ExecutionMemoryPool` below the task's minimum share, and another consumer of the task releases the task's last bytes in the meantime ([SPARK-59444](https://issues.apache.org/jira/browse/SPARK-59444)). Comet's native memory pools reach Spark through `CometTaskMemoryManager.acquireMemory`, which calls `TaskMemoryManager.acquireExecutionMemory` without that guard. A native `try_grow` waiting in Spark gets the exception back as a `CometError` and the task fails. The other consumer can be another native plan in the same task or a JVM consumer such as the shuffle allocator. `grow` logs it and treats the request as overcommit. This reproduces with a real `UnifiedMemoryManager`. With a 100 byte off-heap pool, another task holds 90 bytes and the task holds 10 bytes through a second consumer. A native `acquireMemory(20)` waits below the minimum share of 25. When the second consumer releases its 10 bytes, the waiting request wakes to `key not found: 0`. The Spark fix, [apache/spark#58747](https://github.com/apache/spark/pull/58747), is still open against master. Even once it merges, the Spark versions Comet supports will not have it, so Comet needs its own guard as it did in #6403. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
