[ 
https://issues.apache.org/jira/browse/HDDS-16120?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18103890#comment-18103890
 ] 

Chia-Chuan Ho commented on HDDS-16120:
--------------------------------------

Thanks [~smeng] !
The PR is now open: [https://github.com/apache/ozone/pull/10989]
I’d appreciate your review when you have a chance.

> A single MiniOzoneCluster build timeout cascades into whole-suite failures in 
> MiniOzoneClusterProvider
> ------------------------------------------------------------------------------------------------------
>
>                 Key: HDDS-16120
>                 URL: https://issues.apache.org/jira/browse/HDDS-16120
>             Project: Apache Ozone
>          Issue Type: Bug
>            Reporter: Siyao Meng
>            Assignee: Chia-Chuan Ho
>            Priority: Major
>              Labels: pull-request-available
>         Attachments: HDDS-16120.001.patch
>
>
> {{MiniOzoneClusterProvider.createClusters()}} builds clusters on a background 
> thread and hands them to consumers through a small blocking queue. When a 
> build times out ({{waitForClusterToBeReady()}} throws {{TimeoutException}}) 
> or fails with {{IOException}}, the background thread rethrows it as an 
> unchecked {{RuntimeException("Unable to build cluster")}}. That exception is 
> uncaught on the create thread, so the thread dies and the queue is never 
> refilled. Every subsequent {{provide()}} then blocks until its own timeout 
> and fails with "Failed to obtain available cluster in time".
> The effect is that a single slow or failed cluster build turns into failures 
> for every remaining test that shares the provider. This was observed on 
> master where one method timed out during cluster startup and the following 
> seven methods in the same class all failed with "Failed to obtain available 
> cluster in time". It also makes any transient startup slowdown 
> disproportionately expensive, since one timeout can push an entire 
> integration split to its 90 minute job limit.
> h3. Proposed fix
> Do not let one build failure kill the create thread. Tear down the partial 
> cluster and continue (retry) so a single slow build is not fatal to the 
> suite, or record the failure and surface the original cause from 
> {{provide()}} while keeping the thread alive. Either way removes the "one 
> timeout becomes many failures" behavior.
> h3. Testing
> Unit test for {{MiniOzoneClusterProvider}} where the builder throws on one 
> build: assert that later {{provide()}} calls still succeed (retry path) or 
> fail fast with the original cause, and that the create thread remains alive.
> h3. Note
> This is an amplifier that bounds the blast radius of any single slow cluster 
> build; it is valuable independently of whatever triggers a slow build. It is 
> not itself the root trigger of the recent master integration-job timeouts, 
> and that root trigger is currently unconfirmed.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to