github-actions[bot] commented on code in PR #68610:
URL: https://github.com/apache/doris/pull/68610#discussion_r4129581696
##########
be/src/common/config.cpp:
##########
@@ -335,9 +335,9 @@ DEFINE_Int32(task_executor_max_concurrency_per_task, "-1");
DEFINE_Int32(task_executor_initial_max_concurrency_per_task, "-1");
// Enable task executor in internal table scan.
-DEFINE_Bool(enable_task_executor_in_internal_table, "true");
+DEFINE_Bool(enable_task_executor_in_internal_table, "false");
// Enable task executor in external table scan.
-DEFINE_Bool(enable_task_executor_in_external_table, "true");
+DEFINE_Bool(enable_task_executor_in_external_table, "false");
Review Comment:
[P1] Keep failed remote ThreadPool submissions from running later. This
default selects `ThreadPoolSimplifiedScanScheduler` for newly created workload
groups. If its minimum-worker creation fails, `ThreadPool::init()` still
succeeds with zero workers; `do_submit()` then enqueues a callback before a
failed first-worker attempt returns an error, leaving it queued.
`multiget_data_v2` invokes `done` on that error, but a later successful
submission can run the old callback with released `request`/`response`/`done`
and complete the RPC twice. Make failed submission guarantee that the runnable
cannot execute, and cover the failure/recovery case.
##########
be/src/exec/scan/scanner_context.cpp:
##########
@@ -603,6 +603,14 @@ bool ScannerContext::can_admit_scan_task(const
std::unique_lock<std::mutex>& tra
if (done() || _pending_tasks.empty()) {
return false;
}
+ // Same rule as _pull_next_scan_task() on the TaskExecutor path: once the
shared LIMIT is
+ // exhausted, pending scanners would only open and immediately report EOS,
so do not admit
+ // them while a completed or in-flight task can still wake the operator.
If neither exists,
+ // admit one so it can report EOS and wake the pipeline task.
+ if (_is_shared_scan_limit_exhausted() &&
Review Comment:
[P1] Flush counted projection rows before blocking pending scanners. With
parallel scanning enabled (for example,
`enable_adaptive_pipeline_task_serial_read_on_limit=false`), a projected scan
with LIMIT 4096 can have scanner A buffer 10 rows in `_padding_block` then emit
2500; `Scanner::get_block()` already deducted all 2510 rows from the shared
counter. If scanner B emits 1586 and exhausts the counter, this guard prevents
A from running again while B is completed. After B is consumed,
`get_block_from_queue()` declares EOS, so the query returns only 4086 rows. The
previous ThreadPool path could re-admit A and flush the held 10. Account for
LIMIT on emitted rows or flush counted buffered rows before EOS, and test this
parallel projected case.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]