morningman commented on PR #68713: URL: https://github.com/apache/doris/pull/68713#issuecomment-6082410250
@Gabriel39 here are the end-to-end numbers, and a correction to how I described the trigger. **Correction.** I wrote that a holder sent back to `_pending_tasks` is not reconsidered until a waiter of its context completes. That is too strong. `_pending_tasks` is a stack, so the parked holder is the first one pulled back the next time its context schedules with a slot to spare: when another of its scanners reaches EOS, or when the margin is 2 or more. A context with a single task in flight always resubmits it (`margin_1 = 1`). So the blocking wait recovers as long as some holder keeps running. It stalls only once the threads blocked at the gate keep active + queued at `min_active_file_scan_threads` by themselves. From then on: - every holder whose context has another task in flight is parked when its block is consumed; - a context left with only waiters never schedules again; - the waiters wait for the shares the parked holders keep. That makes "nine concurrent `SELECT *` on 16 cores" wrong. At default settings the stall needs about `min_active_file_scan_threads` readers blocked at once (128 on 16 cores). Your original case, a workload group with fewer remote scan threads than waiting readers, also stalls. The commit message of cddb40bec0b repeats the same overstatement. **Repro setup.** - One Release BE: 18 cores, `-Xmx2g`. - The table: a fluss primary-key table with 16 buckets, 1M of its 2M keys updated after the last kv snapshot. Each bucket declares about 137 MB. - Each query: `SELECT *` with `enable_jni_heap_admission` set and adaptive scan off. - "Waiting on a worker" is this PR at 99e33c47bd1. "Parked" adds cddb40bec0b. Both are on the same master. Wall time per query: | be.conf | queries | waiting on a worker | parked | |---|---|---|---| | `min_active_file_scan_threads = 8` | 1 | 8.5, 9.1 s | 7.3-8.8 s | | same | 2 or 4 at once | 9.4-16.9 s | 9.0-13.2 s | | same | 2 or 4, started 1.5 s apart | 9.5-13.0 s | 7.9-10.2 s | | plus `jni_scanner_heap_budget_ratio = 0.1` (205 MB, one reader at a time) | 2 at once | 127.6 and 131.4 s; next round 214.4 and 222.1 s, both failed with JVM OOM | 11.8-17.3 s over two rounds | With the default budget about seven readers fit at once. So the eight slots hold at most a waiter or two, and the old build always recovered. That is also why none of the A/B reads in the description waited out the limit. With one reader at a time: 1. The first query's eight slots are one holder and seven waiters. 2. The second query's first task is the eighth waiter. 3. The holder is parked as soon as its first block is consumed. 4. At 60 s the old BE logs `A JNI scanner opens after waiting 60003 ms ...: it declared 137 MB, while 1 scanners hold 137 MB of a 204 MB budget and 7 others wait`. 5. The waiters then open together above the budget, and the next wave stalls the same way. In the second round the bursts reached `12 scanners hold 1649 MB of a 204 MB budget`, and the JVM ran out of heap. The parked build never waited past the limit. The fluss regression directory (18 suites) passes on it. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
