github-actions[bot] commented on code in PR #68610:
URL: https://github.com/apache/doris/pull/68610#discussion_r4226650611
##########
be/src/exec/scan/scanner.cpp:
##########
@@ -92,7 +92,10 @@ Status Scanner::get_block_after_projects(RuntimeState*
state, Block* block, bool
const auto& row_descriptor =
_local_state->_parent->operator_row_desc_before_projection();
if (_has_projection) {
_origin_block.clear_column_data(row_descriptor.num_slots());
- if (!_can_merge_padding_blocks(_padding_block, _origin_block)) {
+ // get_block() charges the shared LIMIT as soon as rows pass the
filters. Once peer scanners
+ // exhaust it, the context may finish without running this scanner
again, so rows held in
+ // _padding_block would be charged but never returned. Do not pad
under a shared LIMIT.
+ if (_shared_scan_limit || !_can_merge_padding_blocks(_padding_block,
_origin_block)) {
DORIS_CHECK(_padding_block.empty())
Review Comment:
[P2] Preserve batching for large projected LIMIT scans. This condition
disables padding for every finite shared LIMIT, so with batch size 4096 and one
surviving row per source block, `get_block()` returns that row and each
`ScanTask` publishes it alone. A `LIMIT 10000` scan can then require about
10000 Context submissions and operator wakeups instead of about five half-batch
handoffs when the byte limit permits coalescing. Keep bounded coalescing for
large LIMITs while ensuring only emitted rows are charged and small LIMITs
still make progress; add a selective large-LIMIT throughput case.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]