David Manning created HBASE-30445:
-------------------------------------
Summary: (branch-2) RegionStates may has some expired serverinfo
and make regions do not balance.
Key: HBASE-30445
URL: https://issues.apache.org/jira/browse/HBASE-30445
Project: HBase
Issue Type: Bug
Components: Balancer
Affects Versions: 2.5.4, 2.6.0, 2.7.0
Reporter: David Manning
Assignee: David Manning
It is possible for the balancer to get into a state where it indefinitely seeks
balance but cannot obtain it. LoadBalancer logs will show
{{balancer.StochasticLoadBalancer - Running balancer because cluster has sloppy
server(s).}}
This happens indefinitely because theĀ {{BalancerClusterState}} construction
adds a key twice for the same server. From here, the balancer thinks there is
always a server that has 0 regions, which is "sloppy", and forces a balancer
run even if things are otherwise balanced. The logs may look like below (notice
server 23 is listed twice):
{{balancer.BalancerClusterState - server 24 is on rack 0}}
{{balancer.BalancerClusterState - server 23 is on rack 0}}
{{balancer.BalancerClusterState - server 23 is on rack 0}}
{{balancer.BalancerClusterState - server 22 is on rack 0}}
{{balancer.BalancerClusterState - server 21 is on rack 0}}
The workaround is a restart of hmaster to clear this buggy state.
The likely cause in this deployment, though not verified, was missing
HBASE-29323. As a result, RegionServerReport RPCs could have been significantly
delayed in the queue behind Move RPCs from the RegionMover, causing stale
servers to reappear in the RegionStates.serverMap even after their
ServerCrashProcedure was complete.
However, after investigating the issue, 3.0 branch has deviated from 2.x branch
in a way that appears unintentional. The balancer should provide its baseline
of available servers fromĀ {{onlineServers}} instead of
{{{}serverMap.keySet(){}}}.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)