morningman commented on PR #68713:
URL: https://github.com/apache/doris/pull/68713#issuecomment-6082410250

   @Gabriel39 here are the end-to-end numbers, and a correction to how I 
described the trigger.
   
   **Correction.** I wrote that a holder sent back to `_pending_tasks` is not 
reconsidered until a waiter of its context completes. That is too strong. 
`_pending_tasks` is a stack, so the parked holder is the first one pulled back 
the next time its context schedules with a slot to spare: when another of its 
scanners reaches EOS, or when the margin is 2 or more. A context with a single 
task in flight always resubmits it (`margin_1 = 1`). So the blocking wait 
recovers as long as some holder keeps running.
   
   It stalls only once the threads blocked at the gate keep active + queued at 
`min_active_file_scan_threads` by themselves. From then on:
   - every holder whose context has another task in flight is parked when its 
block is consumed;
   - a context left with only waiters never schedules again;
   - the waiters wait for the shares the parked holders keep.
   
   That makes "nine concurrent `SELECT *` on 16 cores" wrong. At default 
settings the stall needs about `min_active_file_scan_threads` readers blocked 
at once (128 on 16 cores). Your original case, a workload group with fewer 
remote scan threads than waiting readers, also stalls. The commit message of 
cddb40bec0b repeats the same overstatement.
   
   **Repro setup.**
   - One Release BE: 18 cores, `-Xmx2g`.
   - The table: a fluss primary-key table with 16 buckets, 1M of its 2M keys 
updated after the last kv snapshot. Each bucket declares about 137 MB.
   - Each query: `SELECT *` with `enable_jni_heap_admission` set and adaptive 
scan off.
   - "Waiting on a worker" is this PR at 99e33c47bd1. "Parked" adds 
cddb40bec0b. Both are on the same master.
   
   Wall time per query:
   
   | be.conf | queries | waiting on a worker | parked |
   |---|---|---|---|
   | `min_active_file_scan_threads = 8` | 1 | 8.5, 9.1 s | 7.3-8.8 s |
   | same | 2 or 4 at once | 9.4-16.9 s | 9.0-13.2 s |
   | same | 2 or 4, started 1.5 s apart | 9.5-13.0 s | 7.9-10.2 s |
   | plus `jni_scanner_heap_budget_ratio = 0.1` (205 MB, one reader at a time) 
| 2 at once | 127.6 and 131.4 s; next round 214.4 and 222.1 s, both failed with 
JVM OOM | 11.8-17.3 s over two rounds |
   
   With the default budget about seven readers fit at once. So the eight slots 
hold at most a waiter or two, and the old build always recovered. That is also 
why none of the A/B reads in the description waited out the limit.
   
   With one reader at a time:
   1. The first query's eight slots are one holder and seven waiters.
   2. The second query's first task is the eighth waiter.
   3. The holder is parked as soon as its first block is consumed.
   4. At 60 s the old BE logs `A JNI scanner opens after waiting 60003 ms ...: 
it declared 137 MB, while 1 scanners hold 137 MB of a 204 MB budget and 7 
others wait`.
   5. The waiters then open together above the budget, and the next wave stalls 
the same way.
   
   In the second round the bursts reached `12 scanners hold 1649 MB of a 204 MB 
budget`, and the JVM ran out of heap. The parked build never waited past the 
limit. The fluss regression directory (18 suites) passes on it.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to