15767714253 commented on issue #68771:
URL: https://github.com/apache/doris/issues/68771#issuecomment-6094680572

   **Key paradox — the queue is a phantom *and* it grows while the pool is 
idle.**
   
   #### Sampling / 采样(primary evidence)
   
   Consecutive scrapes of `/metrics` on `172.31.17.165` (`ls_normal`, 
`workload_group=normal`) — **10 samples, 10 s apart, 2026-10-10 14:16:29 → 
14:17:59 (GMT+8)**:
   
   ```bash
   for i in $(seq 1 10); do
     curl -s http://172.31.17.165:8040/metrics \
       | grep -E 
'thread_pool_(queue_size|active_threads|submit_failed|task_execution_count_total)\{[^}]*thread_pool_name="ls_normal"'
     sleep 10
   done
   ```
   
   | time (GMT+8) | `queue_size` | `active_threads` | 
`task_execution_count_total` | `submit_failed` |
   | ------------ | ------------ | ---------------- | 
---------------------------- | --------------- |
   | 14:16:29     | 25,610       | 0                | 181,602,147               
   | 0               |
   | 14:16:39     | 25,610       | 0                | 181,606,709               
   | 0               |
   | 14:16:49     | 25,611       | 1                | 181,611,036               
   | 0               |
   | 14:16:59     | 25,611       | 1                | 181,615,106               
   | 0               |
   | 14:17:09     | 25,612       | 0                | 181,618,385               
   | 0               |
   | 14:17:19     | 25,613       | 0                | 181,622,243               
   | 0               |
   | 14:17:29     | 25,613       | 0                | 181,625,538               
   | 0               |
   | 14:17:39     | 25,614       | 0                | 181,630,293               
   | 0               |
   | 14:17:49     | 25,615       | 0                | 181,634,608               
   | 0               |
   | 14:17:59     | 25,615       | 0                | 181,639,119               
   | 0               |
   
   **Derived:**
   
   - `queue_size`: **+5 / 90 s ⇒ ≈ 3.3 / min ⇒ ≈ 4,800 / day**.
   - `task_execution_count_total`: **+36,972 / 90 s ⇒ ≈ 411 splits/s** actually 
consumed by the pool.
   - `max_threads = 48`; `active_threads ∈ {0, 1}` for the whole window; 
`submit_failed = 0` throughout.
   
   **Why this proves the counter is wrong (independent of the "two gauges are 
read under separate locks" caveat):**
   
   The pool reports it is consuming **≈411 splits/s**. If a **25,610-entry real 
queue** were feeding those ≤48 workers, even a single active worker would drain 
it in far under a minute; with **≥47 workers idle at every sample**, a real 
queue of 25,610 would be **emptied in well under 65 s**. Yet across the entire 
90 s window the value **does not fall by a single unit — it only rises**. A 
*real* backlog cannot simultaneously be huge, almost-idle-drained (411/s), and 
monotonically growing. **⇒ The reported depth cannot correspond to real queued 
splits; the gauge is over-counted (phantom queue depth).**
   
   This argument does **not** rely on the atomicity of reading `active_threads` 
and `queue_size` separately — it only relies on the fact that a *real* queue of 
that size cannot survive an almost-idle worker set that is draining the pool at 
~411 splits/s.
   
   #### Earlier corroborating sample / 早期佐证采样(2026-10-08,间隔 4 分钟两次读数)
   
   | BE                     | `queue_size`              | 
`task_execution_count_total`        | `active_threads` |
   | ---------------------- | ------------------------- | 
----------------------------------- | ---------------- |
   | **172.31.20.48**       | 14,893 → **14,905** (+12) | 102,235,465 → 
102,275,040 (+39,575) | 1 → 0            |
   | **172.31.17.165**      | 14,750 → **14,764** (+14) | 106,108,764 → 
106,152,122 (+43,358) | 0 → 0            |
   | 172.31.17.66 (healthy) | 19,030 → **19,030** (+0)  | 3,361,977,363 → 
3,362,017,498       | 0 → 0            |
   
   Rate ≈ **3–3.5 / min ≈ 4,700–5,100 / day**. From the **2026-10-05 14:16 
restart** to the **2026-10-10 14:16 sample (5.00 days)** this predicts ~25,600 
— matching the observed **25,610** almost exactly. Extrapolating, the pool will 
hit the **102400** capacity again around **2026-10-25**, at which point 
`submit_failed` should become non-zero.
   
   Note that a **healthy** BE processes the same volume of tasks (≈40,000 in 4 
minutes) while its `queue_size` stays **perfectly flat at the same ~19,000 
floor** — so this is a **state-dependent accounting leak**, not a 
throughput/backlog problem.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to