[
https://issues.apache.org/jira/browse/HDDS-11128?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Siyao Meng updated HDDS-11128:
------------------------------
Description:
{code}
Tests run: 3, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 102.149 s <<<
FAILURE! - in org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI
org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI.testNodesInDecommissionOrMaintenance(HddsProtos$NodeOperationalState,HddsProtos$NodeOperationalState,
boolean) Time elapsed: 36.908 s <<< ERROR!
TimeoutException:
...
at
org.apache.ozone.test.GenericTestUtils.waitFor(GenericTestUtils.java:204)
at
org.apache.hadoop.ozone.container.TestHelper.waitForReplicaCount(TestHelper.java:444)
at
org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI.testNodesInDecommissionOrMaintenance(TestReconAndAdminContainerCLI.java:264)
{code}
*
https://github.com/adoroszlai/ozone-build-results/blob/master/2024/05/31/31535/it-recon/hadoop-ozone/integration-test/org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI.txt
*
https://github.com/adoroszlai/ozone-build-results/blob/master/2024/07/01/32031/it-recon/hadoop-ozone/integration-test/org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI.txt
h3. Root cause
{{testNodesInDecommissionOrMaintenance}} (decommission case) times out at
{{OzoneTestHelper.waitForReplicaCount}} waiting for a decommission-triggered
replica copy (3 -> 4 for the first node, 4 -> 5 for the second) to be reflected
in SCM. The shared {{waitForReplicaCount}} helper used a fixed 30s budget for
all callers. HDDS-10582 only lowered its poll interval (1000ms -> 200ms) and
kept the 30s total, so on a loaded CI runner the replica copy is not always
observed in time and the test flakes.
h3. Fix
Add a 4-arg {{waitForReplicaCount(containerID, count, cluster, timeoutMillis)}}
overload; the existing 3-arg method delegates with the same 30s default (no
behavior change for the other callers). The decommission/maintenance
replica-copy waits in {{TestReconAndAdminContainerCLI}} now use a 120s budget.
The test stays {{@Flaky("HDDS-11128")}} since a larger timeout cannot prove
non-flakiness.
* https://github.com/apache/ozone/actions/runs/31658391765/job/94319911084
was:
{code}
Tests run: 3, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 102.149 s <<<
FAILURE! - in org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI
org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI.testNodesInDecommissionOrMaintenance(HddsProtos$NodeOperationalState,HddsProtos$NodeOperationalState,
boolean) Time elapsed: 36.908 s <<< ERROR!
TimeoutException:
...
at
org.apache.ozone.test.GenericTestUtils.waitFor(GenericTestUtils.java:204)
at
org.apache.hadoop.ozone.container.TestHelper.waitForReplicaCount(TestHelper.java:444)
at
org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI.testNodesInDecommissionOrMaintenance(TestReconAndAdminContainerCLI.java:264)
{code}
*
https://github.com/adoroszlai/ozone-build-results/blob/master/2024/05/31/31535/it-recon/hadoop-ozone/integration-test/org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI.txt
*
https://github.com/adoroszlai/ozone-build-results/blob/master/2024/07/01/32031/it-recon/hadoop-ozone/integration-test/org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI.txt
> Intermittent timeout during waitForReplicaCount in
> TestReconAndAdminContainerCLI
> --------------------------------------------------------------------------------
>
> Key: HDDS-11128
> URL: https://issues.apache.org/jira/browse/HDDS-11128
> Project: Apache Ozone
> Issue Type: Sub-task
> Components: test
> Reporter: Attila Doroszlai
> Assignee: Siyao Meng
> Priority: Minor
> Labels: pull-request-available
>
> {code}
> Tests run: 3, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 102.149 s <<<
> FAILURE! - in org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI
> org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI.testNodesInDecommissionOrMaintenance(HddsProtos$NodeOperationalState,HddsProtos$NodeOperationalState,
> boolean) Time elapsed: 36.908 s <<< ERROR!
> TimeoutException:
> ...
> at
> org.apache.ozone.test.GenericTestUtils.waitFor(GenericTestUtils.java:204)
> at
> org.apache.hadoop.ozone.container.TestHelper.waitForReplicaCount(TestHelper.java:444)
> at
> org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI.testNodesInDecommissionOrMaintenance(TestReconAndAdminContainerCLI.java:264)
> {code}
> *
> https://github.com/adoroszlai/ozone-build-results/blob/master/2024/05/31/31535/it-recon/hadoop-ozone/integration-test/org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI.txt
> *
> https://github.com/adoroszlai/ozone-build-results/blob/master/2024/07/01/32031/it-recon/hadoop-ozone/integration-test/org.apache.hadoop.ozone.recon.TestReconAndAdminContainerCLI.txt
> h3. Root cause
> {{testNodesInDecommissionOrMaintenance}} (decommission case) times out at
> {{OzoneTestHelper.waitForReplicaCount}} waiting for a decommission-triggered
> replica copy (3 -> 4 for the first node, 4 -> 5 for the second) to be
> reflected in SCM. The shared {{waitForReplicaCount}} helper used a fixed 30s
> budget for all callers. HDDS-10582 only lowered its poll interval (1000ms ->
> 200ms) and kept the 30s total, so on a loaded CI runner the replica copy is
> not always observed in time and the test flakes.
> h3. Fix
> Add a 4-arg {{waitForReplicaCount(containerID, count, cluster,
> timeoutMillis)}} overload; the existing 3-arg method delegates with the same
> 30s default (no behavior change for the other callers). The
> decommission/maintenance replica-copy waits in
> {{TestReconAndAdminContainerCLI}} now use a 120s budget. The test stays
> {{@Flaky("HDDS-11128")}} since a larger timeout cannot prove non-flakiness.
> * https://github.com/apache/ozone/actions/runs/31658391765/job/94319911084
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]