15767714253 commented on issue #68771:
URL: https://github.com/apache/doris/issues/68771#issuecomment-6094680572
**Key paradox — the queue is a phantom *and* it grows while the pool is
idle.**
#### Sampling / 采样(primary evidence)
Consecutive scrapes of `/metrics` on `172.31.17.165` (`ls_normal`,
`workload_group=normal`) — **10 samples, 10 s apart, 2026-10-10 14:16:29 →
14:17:59 (GMT+8)**:
```bash
for i in $(seq 1 10); do
curl -s http://172.31.17.165:8040/metrics \
| grep -E
'thread_pool_(queue_size|active_threads|submit_failed|task_execution_count_total)\{[^}]*thread_pool_name="ls_normal"'
sleep 10
done
```
| time (GMT+8) | `queue_size` | `active_threads` |
`task_execution_count_total` | `submit_failed` |
| ------------ | ------------ | ---------------- |
---------------------------- | --------------- |
| 14:16:29 | 25,610 | 0 | 181,602,147
| 0 |
| 14:16:39 | 25,610 | 0 | 181,606,709
| 0 |
| 14:16:49 | 25,611 | 1 | 181,611,036
| 0 |
| 14:16:59 | 25,611 | 1 | 181,615,106
| 0 |
| 14:17:09 | 25,612 | 0 | 181,618,385
| 0 |
| 14:17:19 | 25,613 | 0 | 181,622,243
| 0 |
| 14:17:29 | 25,613 | 0 | 181,625,538
| 0 |
| 14:17:39 | 25,614 | 0 | 181,630,293
| 0 |
| 14:17:49 | 25,615 | 0 | 181,634,608
| 0 |
| 14:17:59 | 25,615 | 0 | 181,639,119
| 0 |
**Derived:**
- `queue_size`: **+5 / 90 s ⇒ ≈ 3.3 / min ⇒ ≈ 4,800 / day**.
- `task_execution_count_total`: **+36,972 / 90 s ⇒ ≈ 411 splits/s** actually
consumed by the pool.
- `max_threads = 48`; `active_threads ∈ {0, 1}` for the whole window;
`submit_failed = 0` throughout.
**Why this proves the counter is wrong (independent of the "two gauges are
read under separate locks" caveat):**
The pool reports it is consuming **≈411 splits/s**. If a **25,610-entry real
queue** were feeding those ≤48 workers, even a single active worker would drain
it in far under a minute; with **≥47 workers idle at every sample**, a real
queue of 25,610 would be **emptied in well under 65 s**. Yet across the entire
90 s window the value **does not fall by a single unit — it only rises**. A
*real* backlog cannot simultaneously be huge, almost-idle-drained (411/s), and
monotonically growing. **⇒ The reported depth cannot correspond to real queued
splits; the gauge is over-counted (phantom queue depth).**
This argument does **not** rely on the atomicity of reading `active_threads`
and `queue_size` separately — it only relies on the fact that a *real* queue of
that size cannot survive an almost-idle worker set that is draining the pool at
~411 splits/s.
#### Earlier corroborating sample / 早期佐证采样(2026-10-08,间隔 4 分钟两次读数)
| BE | `queue_size` |
`task_execution_count_total` | `active_threads` |
| ---------------------- | ------------------------- |
----------------------------------- | ---------------- |
| **172.31.20.48** | 14,893 → **14,905** (+12) | 102,235,465 →
102,275,040 (+39,575) | 1 → 0 |
| **172.31.17.165** | 14,750 → **14,764** (+14) | 106,108,764 →
106,152,122 (+43,358) | 0 → 0 |
| 172.31.17.66 (healthy) | 19,030 → **19,030** (+0) | 3,361,977,363 →
3,362,017,498 | 0 → 0 |
Rate ≈ **3–3.5 / min ≈ 4,700–5,100 / day**. From the **2026-10-05 14:16
restart** to the **2026-10-10 14:16 sample (5.00 days)** this predicts ~25,600
— matching the observed **25,610** almost exactly. Extrapolating, the pool will
hit the **102400** capacity again around **2026-10-25**, at which point
`submit_failed` should become non-zero.
Note that a **healthy** BE processes the same volume of tasks (≈40,000 in 4
minutes) while its `queue_size` stays **perfectly flat at the same ~19,000
floor** — so this is a **state-dependent accounting leak**, not a
throughput/backlog problem.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]