github-actions[bot] commented on code in PR #68610:
URL: https://github.com/apache/doris/pull/68610#discussion_r4129581696


##########
be/src/common/config.cpp:
##########
@@ -335,9 +335,9 @@ DEFINE_Int32(task_executor_max_concurrency_per_task, "-1");
 DEFINE_Int32(task_executor_initial_max_concurrency_per_task, "-1");
 
 // Enable task executor in internal table scan.
-DEFINE_Bool(enable_task_executor_in_internal_table, "true");
+DEFINE_Bool(enable_task_executor_in_internal_table, "false");
 // Enable task executor in external table scan.
-DEFINE_Bool(enable_task_executor_in_external_table, "true");
+DEFINE_Bool(enable_task_executor_in_external_table, "false");

Review Comment:
   [P1] Keep failed remote ThreadPool submissions from running later. This 
default selects `ThreadPoolSimplifiedScanScheduler` for newly created workload 
groups. If its minimum-worker creation fails, `ThreadPool::init()` still 
succeeds with zero workers; `do_submit()` then enqueues a callback before a 
failed first-worker attempt returns an error, leaving it queued. 
`multiget_data_v2` invokes `done` on that error, but a later successful 
submission can run the old callback with released `request`/`response`/`done` 
and complete the RPC twice. Make failed submission guarantee that the runnable 
cannot execute, and cover the failure/recovery case.



##########
be/src/exec/scan/scanner_context.cpp:
##########
@@ -603,6 +603,14 @@ bool ScannerContext::can_admit_scan_task(const 
std::unique_lock<std::mutex>& tra
     if (done() || _pending_tasks.empty()) {
         return false;
     }
+    // Same rule as _pull_next_scan_task() on the TaskExecutor path: once the 
shared LIMIT is
+    // exhausted, pending scanners would only open and immediately report EOS, 
so do not admit
+    // them while a completed or in-flight task can still wake the operator. 
If neither exists,
+    // admit one so it can report EOS and wake the pipeline task.
+    if (_is_shared_scan_limit_exhausted() &&

Review Comment:
   [P1] Flush counted projection rows before blocking pending scanners. With 
parallel scanning enabled (for example, 
`enable_adaptive_pipeline_task_serial_read_on_limit=false`), a projected scan 
with LIMIT 4096 can have scanner A buffer 10 rows in `_padding_block` then emit 
2500; `Scanner::get_block()` already deducted all 2510 rows from the shared 
counter. If scanner B emits 1586 and exhausts the counter, this guard prevents 
A from running again while B is completed. After B is consumed, 
`get_block_from_queue()` declares EOS, so the query returns only 4086 rows. The 
previous ThreadPool path could re-admit A and flush the held 10. Account for 
LIMIT on emitted rows or flush counted buffered rows before EOS, and test this 
parallel projected case.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to