[ 
https://issues.apache.org/jira/browse/HDDS-15990?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18099332#comment-18099332
 ] 

Ivan Andika commented on HDDS-15990:
------------------------------------

cc: [~tejaskriya] 

> Remove setExclusiveManualCompaction from CompactionService
> ----------------------------------------------------------
>
>                 Key: HDDS-15990
>                 URL: https://issues.apache.org/jira/browse/HDDS-15990
>             Project: Apache Ozone
>          Issue Type: Improvement
>            Reporter: Ivan Andika
>            Assignee: Ivan Andika
>            Priority: Major
>
> We have a production incident because of setExclusiveManualCompaction in 
> CompactionService which caused stalled writes in a cluster with very high 
> write throughput.
> {code:java}
> 2026/07/23-04:14:58.344120 140333138568960 [WARN] [/column_family.cc:1006] 
> [keyTable] Stalling writes because we have 20 level-0 files rate 16777216 
> 2026/07/23-04:16:34.252182 140333356672768 [WARN] [/column_family.cc:1006] 
> [keyTable] Stalling writes because we have 20 level-0 files rate 13421772 
> 2026/07/23-04:19:14.180645 140333511866112 [WARN] [/column_family.cc:1006] 
> [keyTable] Stalling writes because we have 20 level-0 files rate 10737417 
> 2026/07/23-04:19:55.099302 140333419591424 [WARN] [/column_family.cc:1006] 
> [keyTable] Stalling writes because we have 20 level-0 files rate 8589933 
> 2026/07/23-04:20:41.084665 140333149058816 [WARN] [/column_family.cc:1006] 
> [keyTable] Stalling writes because we have 20 level-0 files rate 6871946 
> 2026/07/23-04:21:18.201828 140333237139200 [WARN] [/column_family.cc:1006] 
> [keyTable] Stalling writes because we have 20 level-0 files rate 5497556 
> 2026/07/23-04:21:53.008078 140333214070528 [WARN] [/column_family.cc:1006] 
> [keyTable] Stalling writes because we have 20 level-0 files rate 4398044 
> 2026/07/23-04:22:27.731433 140333300053760 [WARN] [/column_family.cc:1006] 
> [keyTable] Stalling writes because we have 20 level-0 files rate 3518435 ... 
> 2026/07/23-04:42:48.683149 140333660759808 [WARN] [/column_family.cc:1006] 
> [keyTable] Stalling writes because we have 20 level-0 files rate 16384 {code}
>  The supposed events leading to the incident
>  * Normally, RocksDB will schedule automatic compaction that will trigger L0 
> -> L1 compactions
>  * However, due to the exclusive manual compaction, while the manual 
> compaction is run on the bottom levels (which can be very large), these L0 -> 
> L1 automatic compactions are not run
>  * The ongoing high writes and flushes keep creating L0 SST files in the 
> keyTable
>  * Since there are no L0 -> L1 compactions, the L0 files keep increasing 
> until it hits the threshold (20 L0 files)
>  * RocksDB detects this threshold and throttle the writes by 20% until it 
> hits the kMinWriteRate (16KB)
>  * The very low write rate causes OM DB write to be essentially stuck
>  * OM cannot serve the OM RPC write requests quickly and causes the OM RPC 
> handler threads to be blocked
>  * Since all RPC handler threads are blocked, the RPC queue increased until 
> it hits the maximum queue limit
>  * Since RPC queue hits the maximum queue limit, it starts rejecting the 
> request
> We should remove setExclusiveManualCompaction so that automatic compactions 
> can still run to compact L0 -> L1 files and prevent write stalls.
> To fix this we need to ascertain a few things
> Q: Why does setExclusiveManualCompaction set in the first place?
> Seems this is suggested in 
> [https://github.com/apache/ozone/pull/8260#discussion_r2040108204] to reduce 
> redundant compactions running concurrently. However, the extent of the 
> overhead is seems to be hypothetical and the option is just put there "just 
> in case". We need to check the extent of this compaction.
> Q: Even if we setExclusiveManualCompaction = false, will there still be a 
> risk of write stall?
> Seems to be yes, but only if the CPU / Disk resources are occupied which slow 
> down the compaction speed. In our system OM is not running in saturation so 
> we should have enough computation power to ensure that compaction is always 
> fast enough.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to