[
https://issues.apache.org/jira/browse/HBASE-30466?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Jose Luis López updated HBASE-30466:
------------------------------------
Summary: TestSyncReplication and replication/WAL replay tests do not shut
down mini clusters (was: TestSyncReplication* and replication/WAL replay tests
do not shut down mini clusters after a failed start or a hidden @AfterAll)
> TestSyncReplication and replication/WAL replay tests do not shut down mini
> clusters
> -----------------------------------------------------------------------------------
>
> Key: HBASE-30466
> URL: https://issues.apache.org/jira/browse/HBASE-30466
> Project: HBase
> Issue Type: Test
> Reporter: Jose Luis López
> Priority: Major
> Labels: pull-request-available
>
> Some tests do not stop their test clusters when they finish. When surefire
> retries a failed test, the retry runs in the same JVM, so it finds the old
> cluster still running and fails straight away with "A mini-cluster is already
> running". One real failure therefore turns into several fake ones.
> We saw this with TestSyncReplicationStandbyKillMaster on apache/hbase#8743.
> The first run hit a known flake (HBASE-30344). The first retry could not
> start its cluster, and the cleanup skipped the shutdown, so the half-started
> cluster was left running. The second retry then failed because of that
> leftover cluster.
> h3. Affected tests
> * The TestSyncReplication* tests: cleanup skips shutting down a cluster that
> failed to start, and one cleanup error stops the remaining clusters from
> being shut down.
> * TestClaimReplicationQueue, TestRemovePeerProcedureWaitForSCP and
> TestAsyncWALReplay (plus its 3 subclasses): their cleanup method replaces the
> parent class's cleanup instead of adding to it, so their clusters are never
> shut down.
> h3. Fix
> Always shut down every cluster, even when starting it or an earlier cleanup
> step failed. Test-only change.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)