dwsmith1983 opened a new issue, #6224: URL: https://github.com/apache/datafusion-comet/issues/6224
### Describe the bug Spark's `ExecutionMemoryPool` removes a task's `memoryForTask` entry when the task's balance reaches zero. An acquire that is parked in `ExecutionMemoryPool.acquireMemory` reads that entry when it wakes up, and throws `java.util.NoSuchElementException: key not found: <taskAttemptId>` if the entry has been removed. After #5613 the `fair_unified` pool holds one anchor byte with Spark, which keeps the entry alive, and none of the pool's own releases can zero the balance. One window is still open. Spark declines the anchor when the task is already at its share. Until a later grow's retry succeeds, the pool can hold nothing from Spark while one of its requests (the anchor retry or the real request) is parked. If another off-heap consumer of the same task then frees its last bytes, the balance reaches zero and the entry is removed while the pool's request is still waiting. The shuffle allocator (`CometUnifiedShuffleMemoryAllocator`) and any Spark operator in the task are examples of such a consumer. When the parked request fails, `try_grow` rolls back its charge and returns the error. The error is not `ResourcesExhausted`, so the operator cannot spill and the task fails. `grow` logs a warning and carries the bytes as overcommit. ### Steps to reproduce This is a component-level sketch in the style of `CometTaskMemoryManagerSuite`. A full query reproduction has not been established. 1. Create a 100-byte off-heap `UnifiedMemoryManager`. Give task 0 a `TaskMemoryManager` and a `CometTaskMemoryManager`, and give it a second off-heap `MemoryConsumer` that stands in for the shuffle allocator. 2. The sibling consumer acquires 100 bytes. `acquireAnchor(1)` returns 0, so the anchor is declined. 3. Another task acquires 90 bytes. The sibling frees 90, which leaves it holding 10 and leaves nothing free. 4. On another thread, `acquireAnchor(1)` (or `acquireMemory(n)`) parks: the task holds 10, which is below the 25-byte minimum share for two active tasks. 5. The sibling frees its last 10 bytes. The entry is removed, and the parked acquire throws `NoSuchElementException`. ### Expected behavior A native acquire that is parked in Spark completes, or is refused with a grant of zero. It does not fail because another consumer of the same task released memory. ### Additional context **Why main has it too.** Main has no anchor at all. Any native acquire parked while the task's balance is made up only of a JVM consumer's bytes fails the same way when that consumer frees them. The native pool's own releases can also zero the balance on main. #5613 closes that case and narrows this one to the period before the anchor is first held. **Option: take the anchor when the pool is created.** Acquiring the byte in the pool constructor, rather than on the first grow, would cover the period before the first grow. It has costs: - Every native plan makes a JNI call into Spark's memory manager at creation, including plans that never allocate natively. - Such plans would count as active tasks in Spark. That lowers the 1/N and 1/(2N) shares of the other tasks on the executor, which is why #5613 takes the anchor lazily. - It does not close the window. Spark can decline the byte at creation for the same reason it does now, when a JVM consumer already holds the task's share. The pool would then be in exactly this state. A complete fix probably needs the release side to be coordinated. One way is to route the JVM consumers' releases through something that knows whether a native acquire is parked. Another is to add a Spark-side hook. Both are larger than an anchor change. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
