Jose Luis López created HBASE-30466:
---------------------------------------
Summary: TestSyncReplication* and replication/WAL replay tests do
not shut down mini clusters after a failed start or a hidden @AfterAll
Key: HBASE-30466
URL: https://issues.apache.org/jira/browse/HBASE-30466
Project: HBase
Issue Type: Test
Reporter: Jose Luis López
On apache/hbase#8743 (large-wave-3), TestSyncReplicationStandbyKillMaster
failed three times in a row, but only the first failure was real:
# Run 1 failed in the test body (the flake tracked in HBASE-30344).
# Rerun 1: {{UTIL1.startMiniCluster}}
({{SyncReplicationTestBaseNoBeforeAll.startClusters:111}}) started DFS, then
the master did not initialise in 200 s, so it threw "IOException: Shutting
down". DFS stayed up and the util stayed marked running.
# {{@AfterAll}} called {{shutdown(UTIL1)}}, which returns early when
{{util.getHBaseCluster() == null}} and so never called
{{shutdownMiniCluster()}}.
# Rerun 2 failed at once with "IllegalStateException: A mini-cluster is already
running"; its log still shows the leaked master writing to the leaked DFS.
h3. Problems
* {{SyncReplicationTestBaseNoBeforeAll.shutdown(util)}} skips
{{shutdownMiniCluster()}} exactly in the state a failed start leaves behind.
* {{tearDown()}} runs admin calls and then shuts down UTIL1, UTIL2 and ZK in
sequence with no finally, so one failing call leaks the rest. This affects all
15 subclasses.
* {{TestClaimReplicationQueue}} and {{TestRemovePeerProcedureWaitForSCP}}
declare a static {{tearDownAfterClass()}} that hides
{{TestReplicationBaseNoBeforeAll.tearDownAfterClass()}} without calling it, so
their two clusters are never shut down.
* {{TestAsyncWALReplay}} does the same with
{{AbstractTestWALReplay.tearDownAfterClass()}}, which affects it and its 3
subclasses.
h3. Fix
Always call {{shutdownMiniCluster()}} in finally, chain UTIL1 -> UTIL2 -> ZK
through finally, and make the hiding methods call the base teardown in finally.
Test-only change.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)