15767714253 opened a new issue, #68771:
URL: https://github.com/apache/doris/issues/68771

   ### Search before asking
   
   - [x] I had searched in the 
[issues](https://github.com/apache/doris/issues?q=is%3Aissue) and found no 
similar issues.
   
   
   ### Version
   
   4.1.2
   
   ### What's Wrong?
   
   The Prometheus metric 
`doris_be_thread_pool_queue_size{thread_pool_name="ls_normal", 
workload_group="normal"}` grows **monotonically** on **2 of 5 BEs** 
(`172.31.20.48`, `172.31.17.165`), starting around **2026-09-17**, until it 
saturates near **102400** (which equals `doris_scanner_thread_pool_queue_size`) 
around **2026-10-05**, then drops vertically.
   
   The vertical drop is **not** the queue draining — it coincides exactly with 
a **BE restart**:
   
   | BE                | `LastStartTime` (SHOW BACKENDS) | Growing?  |
   | ----------------- | ------------------------------- | --------- |
   | 172.31.17.66      | 2026-06-24 14:52:14             | No (flat) |
   | 172.31.25.179     | 2026-06-24 14:50:20             | No (flat) |
   | 172.31.19.197     | 2026-06-24 14:50:48             | No (flat) |
   | **172.31.20.48**  | **2026-10-05 14:17:36**         | **Yes**   |
   | **172.31.17.165** | **2026-10-05 14:16:37**         | **Yes**   |
   
   The two affected BEs also have ~1/33 of the cumulative 
`task_execution_count_total` of the other three, consistent with a recent 
restart.
   
   **Key paradox — the queue is a phantom.** At the time of sampling, the 
affected pools report:
   
   ```
   doris_be_thread_pool_active_threads{thread_pool_name="ls_normal"} = 0      
(max_threads = 48)
   doris_be_thread_pool_submit_failed{thread_pool_name="ls_normal"} = 0
   doris_be_thread_pool_queue_size{thread_pool_name="ls_normal"}    = 14905   
(!!)
   ```
   
   **Idle workers (active = 0) together with a non-empty queue of ~15,000 
entries is impossible for a correct implementation** — the dispatcher would 
drain any real pending split immediately. This indicates the gauge is 
**over-counted (phantom queue depth)**, while the real queue is empty.
   
   **It is still growing right now** (two samples, 4 minutes apart):
   
   | BE                     | `queue_size`            | 
`task_execution_count_total`        | `active_threads` |
   | ---------------------- | ----------------------- | 
----------------------------------- | ---------------- |
   | **172.31.20.48**       | 14893 → **14905** (+12) | 102,235,465 → 
102,275,040 (+39,575) | 1 → 0            |
   | **172.31.17.165**      | 14750 → **14764** (+14) | 106,108,764 → 
106,152,122 (+43,358) | 0 → 0            |
   | 172.31.17.66 (healthy) | 19030 → **19030** (+0)  | 3,361,977,363 → 
3,362,017,498       | 0 → 0            |
   
   Rate ≈ **3–3.5 / minute ≈ 4,700 / day**. From the 2026-10-05 restart to the 
2026-10-08 sample (2.9 days) this accumulates to ~14,900 — matching the 
observed value exactly. Extrapolating, the pool will hit the **102400** 
capacity again around **2026-10-25**, at which point `submit_failed` should 
become non-zero.
   
   Note that a **healthy** BE processing the same volume of tasks (≈40,000 in 4 
minutes) keeps `queue_size` **perfectly flat** — so this is a **state-dependent 
accounting leak**, not a throughput/backlog problem.
   
   
   
   **config:**
   # 基础参数
   # 系统日志文件保留的滚动数量,设置为 1 表示只保留最新的日志文件。
   sys_log_roll_num = 7
   
   # Stream Load 导入 JSON 数据的最大文件大小限制(2048 MB = 2GB)。
   streaming_load_json_max_mb = 2048
   
   # 允许 STRING 类型字段存储的最大字节数的软限制(1 GB),超出可能影响性能。
   string_type_length_soft_limit_bytes = 1073741824
   
   # 合并参数
   # 累计合并(Cumulative Compaction)任务的最大并发线程数。
   max_cumu_compaction_threads = 24
   max_base_compaction_threads = 8
   
   # 每个磁盘同时运行的 Compaction 任务数量上限。
   compaction_task_num_per_disk = 32
   
   # 每个SSD磁盘同时运行的 Compaction 任务数量上限。
   compaction_task_num_per_fast_disk = 32
   
   # 触发 Cumulative Compaction 的最小数据版本增量(Delta)数量阈值。
   cumulative_compaction_min_deltas = 5
   
   # 触发 Base Compaction 的条件
   base_compaction_interval_seconds_since_last_operation = 43200
   base_compaction_min_data_ratio = 0.2
   
   # 用 Segment 层级压缩(节省 CPU 资源,推荐关闭)。
   enable_segcompaction = true
   
   # 列compatction
   enable_vertical_compaction = true
   # 禁用 Compaction 任务的优先级调度机制(减少调度开销)。
   enable_compaction_priority_scheduling = false
   
   # 写入优化
   # 写入缓冲区的内存大小(1 GB),提升高频写场景的吞吐量。
   write_buffer_size = 1073741824
   
   # 线程池
   # Pipeline 执行线程数(用于并行查询处理)。
   pipeline_executor_size = 12
   
   # 阻塞式 Pipeline 执行线程数(处理阻塞 I/O 操作)。
   blocking_pipeline_executor_size = 8
   
   # 数据版本发布任务的并发工作线程数。
   publish_version_worker_count = 16
   
   # 文件上传任务的线程数(影响存算分离上传性能)。
   upload_worker_count = 16
   
   # Tablet 数据写入线程数(提升并发写入能力)。
   number_tablet_writer_threads = 32
   
   # 高优先级数据推送任务的线程数(如副本同步)。
   push_worker_count_high_priority = 6
   
   # 普通优先级数据推送任务的线程数。
   push_worker_count_normal_priority = 6
   
   # 批量数据发送线程池的线程数(影响节点间数据传输)。
   send_batch_thread_pool_thread_num = 96
   
   #默认只有 2 个线程,在高写入量下是瓶颈。可增加该值让更多 CPU 参与刷盘
   flush_thread_num_per_store = 8
   
   
   # S3 存算分离
   # S3 文件上传线程池的最大线程数。
   num_s3_file_upload_thread_pool_max_thread = 64
   
   # S3 文件上传线程池的最小线程数(核心数)。
   num_s3_file_upload_thread_pool_min_thread = 32
   
   # 操作 S3 文件系统的最大线程数。
   max_s3_file_system_thread_num = 128
   
   # 操作 S3 文件系统的最小线程数(核心数)。
   min_s3_file_system_thread_num = 32
   
   remote_storage_read_buffer_mb = 128
   doris_remote_scanner_thread_pool_thread_num = 128
   
   #S3 写入缓冲区大小(字节),需 ≥ 5MB
   s3_write_buffer_size = 10485760
   #S3 文件系统本地上传缓冲区大小
   s3_file_system_local_upload_buffer_size = 10485760
   
   
   # 缓存配置
   #开启行缓存
   disable_storage_row_cache = false 
   
   #行缓存内存比例   
   row_cache_mem_limit = 10%
   #缓存
   storage_page_cache_limit = 15%
   
   max_tablet_version_num = 4000
   
   enable_packed_file = false
   
   # delete bitmap 计算调整
   calc_delete_bitmap_worker_count = 16
   calc_tablet_delete_bitmap_task_max_thread = 64
   delete_bitmap_agg_cache_capacity = 1073741824
   calc_delete_bitmap_max_thread = 64
   calc_delete_bitmap_for_load_max_thread = 16
   tablet_publish_txn_max_thread  = 64
   
   
   ### What You Expected?
   
   1. `doris_be_thread_pool_queue_size` for `ls_normal` should fall back to ~0 
when the pool is idle (`active_threads = 0`), and must never grow without bound.
   2. The counter should stay consistent across token lifecycle events (token 
shutdown / task removal) — every entry removed from the queue must decrement 
`_total_queued_tasks`.
   3. `submit_failed` should not begin incrementing merely because of a 
leaked/phantom queue count.
   
   ### How to Reproduce?
   
   _No response_
   
   ### Anything Else?
   
   _No response_
   
   ### Are you willing to submit PR?
   
   - [ ] Yes I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to