Jose Luis López created HBASE-30460:
---------------------------------------

             Summary: RegionServer abort timer is never cancelled and halts the 
JVM after the RegionServer has shut down
                 Key: HBASE-30460
                 URL: https://issues.apache.org/jira/browse/HBASE-30460
             Project: HBase
          Issue Type: Bug
          Components: regionserver
            Reporter: Jose Luis López


HRegionServer#abort calls scheduleAbortTimer(), which creates the
"Abort regionserver monitor" Timer and schedules SystemExitWhenAbortTimeout
to run after hbase.regionserver.abort.timeout (default 1200000 ms). The task
prints a thread dump headed "Zombie HRegionServer" and calls
Runtime.getRuntime().halt(1).

The timer is never cancelled. It fires even when the abort completes
normally and HRegionServer#run has returned. As introduced by HBASE-21325
("Force to terminate regionserver when abort hang in somewhere") and
switched to halt by HBASE-21932, its purpose is to cap an abort that hangs.
After a clean exit there is nothing left to cap.

In a standalone RegionServer process this goes unnoticed, because the
process exits once run() returns. When the RegionServer runs inside a JVM
that outlives it, as with HBaseTestingUtility / MiniHBaseCluster, the timer
halts that JVM 20 minutes after an abort that had already finished. Neither
HBaseTestingUtility nor MiniHBaseCluster sets
hbase.regionserver.abort.timeout or hbase.regionserver.abort.timeout.task.

Observed downstream in Apache Hadoop, whose
hadoop-yarn-server-timelineservice-hbase-tests module runs a mini-cluster
in the Maven JVM (surefire forkCount=0), with HBase 2.6.3-hadoop3:

  22:52:02.488  main: ***** STOPPING region server '…,42315,…' *****
                (HBaseTestingUtility#shutdownMiniCluster)
  22:52:02.494  RS_OPEN_REGION: Opened f6f5937a47a606f0399f03df7faef64f
  22:52:02.498  AssignRegionHandler: Fatal error occurred while opening
                region …, aborting…
                RegionServerStoppedException: Server … stopping
                  at RSRpcServices.checkOpen(RSRpcServices.java:1571)
                  at HRegionServer.postOpenDeployTasks(HRegionServer.java:2608)
                  at AssignRegionHandler.process(AssignRegionHandler.java:161)
  22:52:02.501  ***** ABORTING region server …,42315,…: Failed to open
                region … and can not recover *****
  22:52:02.825  RS:0;…:42315 Exiting; stopping=…,42315,…; zookeeper
                connection closed.        <- run() returned, shutDown=true
  22:52:02.826  JVMClusterUtil: Shutdown of 1 master(s) and 1
                regionserver(s) complete
  …
  23:12:02.528  Process Thread Dump: Zombie HRegionServer
                -> Runtime.halt(1); the whole Maven build exits with code 1

The RegionServer had exited 20 minutes before it was declared a zombie.

Proposed fix: cancel abortMonitor once the RegionServer has finished
shutting down, e.g. at the end of HRegionServer#run after shutDown is set:

    if (abortMonitor != null) {
      abortMonitor.cancel();
    }

An abort that really hangs never reaches that point, so the timer still
fires in the case it exists for. A test can use
hbase.regionserver.abort.timeout.task to install a task that records
whether it ran. With a short hbase.regionserver.abort.timeout, it can abort
a RegionServer in a mini-cluster, wait for the RegionServer thread to end,
and assert the task never ran.

Related, possibly a separate issue: AssignRegionHandler#handleException
aborts the server when postOpenDeployTasks fails with
RegionServerStoppedException because a requested stop is already under
way. That turns an ordinary shutdown racing a region open into an abort,
which is what armed the timer here.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to