github-actions[bot] commented on code in PR #67845:
URL: https://github.com/apache/doris/pull/67845#discussion_r4226760561


##########
be/src/runtime/workload_group/workload_group_manager.cpp:
##########
@@ -823,40 +848,93 @@ bool WorkloadGroupMgr::handle_single_query_(const 
std::shared_ptr<ResourceContex
             requestor->task_controller()->cancel(error_status);
             return true;
         }
-    } else {
-        // Should not consider about process memory. For example, the query's 
limit is 100g, workload
-        // group's memlimit is 10g, process memory is 20g. The query reserve 
will always failed in wg
-        // limit, and process is always have memory, so that it will resume 
and failed reserve again.
-        const size_t test_memory_size = std::max<size_t>(size_to_reserve, 32L 
* 1024 * 1024);
-        if 
(!GlobalMemoryArbitrator::is_exceed_soft_mem_limit(test_memory_size)) {
+    }
+}
+
+void WorkloadGroupMgr::spill_query_(const std::shared_ptr<ResourceContext>& 
requestor) {
+    SCOPED_ATTACH_TASK(requestor);
+    auto status = requestor->task_controller()->revoke_memory();
+    if (!status.ok()) {
+        requestor->task_controller()->cancel(status);
+    }
+}
+
+bool WorkloadGroupMgr::resolve_process_memory_exceeded_query_(
+        const std::shared_ptr<ResourceContext>& requestor, size_t 
size_to_reserve,
+        int64_t time_in_queue, size_t memory_usage, bool has_running_task) {
+    const auto query_id = print_id(requestor->task_controller()->task_id());
+    const auto wg = requestor->workload_group();
+
+    // The caller observed the pressure before it inspected the pipeline tasks 
and tried to
+    // revoke memory from other workload groups, during which other queries or 
the cache may have
+    // released memory. Re-check the recorded reservation right before acting, 
so that the query
+    // is never cancelled from a stale observation.
+    if (!GlobalMemoryArbitrator::is_exceed_soft_mem_limit(size_to_reserve)) {
+        LOG(INFO) << "Query: " << query_id << ", process limit not exceeded 
now, resume this query"
+                  << ", process memory info: "
+                  << GlobalMemoryArbitrator::process_memory_used_details_str()
+                  << ", wg info: " << wg->debug_string();
+        requestor->task_controller()->set_memory_sufficient(true);
+        return true;
+    }
+
+    const bool exceed_hard_mem_limit = 
GlobalMemoryArbitrator::is_exceed_hard_mem_limit();
+    if (has_running_task) {
+        // Spilling needs every task of the query to be idle. Below the hard 
limit wait for the
+        // running task to yield, the query is handled again in the next round.
+        if (!exceed_hard_mem_limit) {

Review Comment:
   [P2] Apply the timeout even while another task is running. This branch 
returns before the elapsed-time check, so a query paused under persistent soft 
pressure can remain blocked past `spill_in_paused_queue_timeout_ms` for as long 
as a sibling stays in one operator call. The new test backdates past the 
timeout and expects that continued wait, although this PR promises a bounded 
grace period. Cancelling a running query is already supported at the hard 
limit; use the elapsed timeout for that cancellation too, while continuing to 
avoid spilling a running task. This is separate from the existing hard-limit 
thread.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to