[
https://issues.apache.org/jira/browse/SPARK-59461?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-59461:
-----------------------------------
Labels: pull-request-available (was: )
> Reclaim optional execution memory before active query allocations
> -----------------------------------------------------------------
>
> Key: SPARK-59461
> URL: https://issues.apache.org/jira/browse/SPARK-59461
> Project: Spark
> Issue Type: Improvement
> Components: Spark Core
> Affects Versions: 5.0.0
> Reporter: Chao Sun
> Assignee: Chao Sun
> Priority: Major
> Labels: pull-request-available
>
> h2. Motivation
> Readers and plugins can improve throughput by preparing data before it is
> needed, for example by downloading the next selected Parquet row group while
> decoding the current one. Memory holding unused speculative work should have
> lower priority than the memory needed to make progress on the query.
> Admitting speculative buffers only when execution memory is free does not
> protect later ordinary allocations. For example, in a 1,000-byte execution
> pool, a task can reserve 400 bytes for unused read-ahead before asking for
> 700 bytes of ordinary working memory. Leaving the optional reservation in
> place leaves only 600 bytes available, even though discarding the unused
> buffer would allow the ordinary request to succeed. Reservations from another
> task can also affect the requesting task's fair share.
> TaskMemoryManager can ask consumers within its own task to spill, and
> UnifiedMemoryManager can evict borrowed storage memory. Neither mechanism
> gives an ordinary allocation a way to request release of unused optional
> execution buffers belonging to other tasks. Unmanaged-memory reporting
> adjusts effective capacity, but does not supply a release callback or
> reservation protocol.
> This is a memory-management enhancement to support cooperating
> readers/plugins, not a claim that an existing built-in Spark reader has a
> correctness bug.
> h2. Proposed scope
> Add the internal executor memory policy for optional, task-attributed
> execution-memory reservations. Keep public TaskMemoryManager APIs,
> per-consumer/task-lifecycle integration, and reader-specific prefetch
> scheduling in separate follow-ups. This first change does not enable
> read-ahead, add a configuration, or change file selection or page skipping.
> h2. Admission and reclamation
> * Optional admission grants the full request from currently free execution
> memory or returns zero without changing task registration or retaining a
> partial reservation. It does not borrow storage capacity, evict blocks,
> spill, or wait for capacity. Short bookkeeping locks may still contend; the
> API is not wait-free.
> * Owners register release-only callbacks before optional admission. Ordinary
> allocation first checks whether its full request can be granted using the
> existing allocation policy without eviction or a capacity wait. This check
> must include the prospective task fair share, not just aggregate free bytes.
> * When the request fits, retain the optional buffers. Otherwise invoke
> matching owners across the executor, then retry through Spark's existing
> allocator. Reclamation does not promise that every ordinary request can be
> granted in full.
> * A shared admission gate prevents fresh optional reservations from racing
> ordinary reclamation. Invoke callbacks outside both the memory-manager
> monitor and the registration monitor.
> * MemoryStore operations that take the manager monitor before making nested
> allocations must establish the reclamation boundary before taking that
> monitor. The first implementation conservatively drains optional owners at
> these outer boundaries; avoiding unnecessary cancellation there is a separate
> refinement.
> * Account for current upstream unmanaged-memory usage in new admission and
> pressure preflights, consistently with ordinary admission.
> {noformat}
> Optional request --> free capacity and fair share available?
> yes: reserve all bytes
> no: return zero; use ordinary demand work
> Ordinary request --> full immediate grant possible?
> yes: keep optional buffers; grant normally
> no: release unused optional work
> retry existing allocation policy
> {noformat}
> h2. Ownership and failure contract
> Callbacks may be concurrent, repeated, or already captured when an owner
> unregisters. Each owner must release a reservation exactly once, drain before
> unregistering, and tolerate later calls. A callback must not allocate
> execution memory, acquire a task-manager monitor, wait for I/O/task cleanup,
> or destroy a whole reader/session. Callers must not hold a lock needed by a
> reclaimer while requesting ordinary memory.
> Failed reclamation must not manufacture free memory or silently remove the
> owner. Keep its charge and registration available for retry. Complete
> release-only block removal/clearing even when an optional callback throws;
> otherwise BlockManager metadata and MemoryStore contents can diverge. Failed
> unrolling or final storage transfer must dispose consumed resources and
> release only that operation's unroll credit, preserving the original error
> and normal partial-result ownership.
> h2. Related work and non-goals
> SPARK-56918 proposes shrinking executor-wide external storage caches (for
> example, a native columnar cache). This proposal instead concerns
> task-attributed, disposable execution buffers with all-or-zero admission. The
> two efforts should coordinate allocator/lock contracts, but this issue does
> not implement the external-cache SPI or move its accounting to the execution
> pool.
> Do not replace ordinary spilling, redefine Spark's general fairness policy,
> reclaim buffers already in use, or claim an automatic speedup for workloads
> without an adopting consumer. Re-fetching discarded data and added allocator
> bookkeeping are performance risks, not free operations.
> h2. Validation
> Cover full/denied optional admission, no-owner and ample-capacity ordinary
> paths, same-task and cross-task pressure, on/off-heap isolation, callback
> lock ordering and failure handling, storage/unroll cleanup, and upstream
> unmanaged-memory accounting. Use deterministic concurrency controls rather
> than timing assumptions. Validate through real MemoryStore/BlockManager paths
> where ownership matters. Measure no-owner allocation overhead separately from
> consumer-level end-to-end benefit; do not infer either from passing
> correctness tests.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]