Jose Luis López created HBASE-30460:
---------------------------------------
Summary: RegionServer abort timer is never cancelled and halts the
JVM after the RegionServer has shut down
Key: HBASE-30460
URL: https://issues.apache.org/jira/browse/HBASE-30460
Project: HBase
Issue Type: Bug
Components: regionserver
Reporter: Jose Luis López
HRegionServer#abort calls scheduleAbortTimer(), which creates the
"Abort regionserver monitor" Timer and schedules SystemExitWhenAbortTimeout
to run after hbase.regionserver.abort.timeout (default 1200000 ms). The task
prints a thread dump headed "Zombie HRegionServer" and calls
Runtime.getRuntime().halt(1).
The timer is never cancelled. It fires even when the abort completes
normally and HRegionServer#run has returned. As introduced by HBASE-21325
("Force to terminate regionserver when abort hang in somewhere") and
switched to halt by HBASE-21932, its purpose is to cap an abort that hangs.
After a clean exit there is nothing left to cap.
In a standalone RegionServer process this goes unnoticed, because the
process exits once run() returns. When the RegionServer runs inside a JVM
that outlives it, as with HBaseTestingUtility / MiniHBaseCluster, the timer
halts that JVM 20 minutes after an abort that had already finished. Neither
HBaseTestingUtility nor MiniHBaseCluster sets
hbase.regionserver.abort.timeout or hbase.regionserver.abort.timeout.task.
Observed downstream in Apache Hadoop, whose
hadoop-yarn-server-timelineservice-hbase-tests module runs a mini-cluster
in the Maven JVM (surefire forkCount=0), with HBase 2.6.3-hadoop3:
22:52:02.488 main: ***** STOPPING region server '…,42315,…' *****
(HBaseTestingUtility#shutdownMiniCluster)
22:52:02.494 RS_OPEN_REGION: Opened f6f5937a47a606f0399f03df7faef64f
22:52:02.498 AssignRegionHandler: Fatal error occurred while opening
region …, aborting…
RegionServerStoppedException: Server … stopping
at RSRpcServices.checkOpen(RSRpcServices.java:1571)
at HRegionServer.postOpenDeployTasks(HRegionServer.java:2608)
at AssignRegionHandler.process(AssignRegionHandler.java:161)
22:52:02.501 ***** ABORTING region server …,42315,…: Failed to open
region … and can not recover *****
22:52:02.825 RS:0;…:42315 Exiting; stopping=…,42315,…; zookeeper
connection closed. <- run() returned, shutDown=true
22:52:02.826 JVMClusterUtil: Shutdown of 1 master(s) and 1
regionserver(s) complete
…
23:12:02.528 Process Thread Dump: Zombie HRegionServer
-> Runtime.halt(1); the whole Maven build exits with code 1
The RegionServer had exited 20 minutes before it was declared a zombie.
Proposed fix: cancel abortMonitor once the RegionServer has finished
shutting down, e.g. at the end of HRegionServer#run after shutDown is set:
if (abortMonitor != null) {
abortMonitor.cancel();
}
An abort that really hangs never reaches that point, so the timer still
fires in the case it exists for. A test can use
hbase.regionserver.abort.timeout.task to install a task that records
whether it ran. With a short hbase.regionserver.abort.timeout, it can abort
a RegionServer in a mini-cluster, wait for the RegionServer thread to end,
and assert the task never ran.
Related, possibly a separate issue: AssignRegionHandler#handleException
aborts the server when postOpenDeployTasks fails with
RegionServerStoppedException because a requested stop is already under
way. That turns an ordinary shutdown racing a region open into an abort,
which is what armed the timer here.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)