[ 
https://issues.apache.org/jira/browse/HBASE-30445?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

David Manning updated HBASE-30445:
----------------------------------
    Description: 
It is possible for the balancer to get into a state where it indefinitely seeks 
balance but cannot obtain it. LoadBalancer logs will show

{{balancer.StochasticLoadBalancer - Running balancer because cluster has sloppy 
server(s).}}

This happens indefinitely because the {{BalancerClusterState}} construction 
adds a key twice for the same server. From here, the balancer thinks there is 
always a server that has 0 regions, which is "sloppy", and forces a balancer 
run even if things are otherwise balanced. The logs may look like below (notice 
server 23 is listed twice):

{{balancer.BalancerClusterState - server 24 is on rack 0}}

{{balancer.BalancerClusterState - server 23 is on rack 0}}

{{balancer.BalancerClusterState - server 23 is on rack 0}}

{{balancer.BalancerClusterState - server 22 is on rack 0}}

{{balancer.BalancerClusterState - server 21 is on rack 0}}

The workaround is a restart of hmaster to clear this buggy state.

The likely cause in this deployment, though not verified, was missing 
HBASE-29323. As a result, ReportProcedureDone RPCs could have been 
significantly delayed/deadlocked in the queue behind Move RPCs from the 
RegionMover, causing stale servers to reappear in the RegionStates.serverMap 
even (long) after their ServerCrashProcedure was complete. Ultimately, this is 
the same type of cause described by HBASE-23564.

However, after investigating the issue, 3.0 branch has deviated from 2.x branch 
in a way that appears unintentional. The balancer should provide its baseline 
of available servers from {{onlineServers}} instead of 
{{{}serverMap.keySet(){}}}.

3.0 branch history:
 # HBASE-22523 adds as {{{}serverMap.keySet(){}}}: 
[https://github.com/apache/hbase/commit/04e5bf96d84ff2f111b1e879c8daa5360b621922#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R552-R556]
 
 # HBASE-23564 changes to {{{}onlineServers{}}}: 
[https://github.com/apache/hbase/commit/ab40b9648b20cef6cb041789e4fbb11e4d36e3c2#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R562]
 

2.x branch history:
 # HBASE-22523 adds as {{{}serverMap.keySet(){}}}: 
[https://github.com/apache/hbase/commit/14041b39f2da17de7a654a1733f3fa136fe13f63#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R552-R556]
 # HBASE-23564 changes to {{{}onlineServers{}}}: 
[https://github.com/apache/hbase/commit/7a0e4d814059e6a26abc834770e5f66231e0400e#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R563]
 
 # HBASE-23102 changes back to {{{}serverMap.keySet(){}}}: 
[https://github.com/apache/hbase/commit/3b6d8c3394a22c71b9836c4000dbd582af248aa3#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R585]
 

  was:
It is possible for the balancer to get into a state where it indefinitely seeks 
balance but cannot obtain it. LoadBalancer logs will show

{{balancer.StochasticLoadBalancer - Running balancer because cluster has sloppy 
server(s).}}

This happens indefinitely because the {{BalancerClusterState}} construction 
adds a key twice for the same server. From here, the balancer thinks there is 
always a server that has 0 regions, which is "sloppy", and forces a balancer 
run even if things are otherwise balanced. The logs may look like below (notice 
server 23 is listed twice):

{{balancer.BalancerClusterState - server 24 is on rack 0}}

{{balancer.BalancerClusterState - server 23 is on rack 0}}

{{balancer.BalancerClusterState - server 23 is on rack 0}}

{{balancer.BalancerClusterState - server 22 is on rack 0}}

{{balancer.BalancerClusterState - server 21 is on rack 0}}

The workaround is a restart of hmaster to clear this buggy state.

The likely cause in this deployment, though not verified, was missing 
HBASE-29323. As a result, ReportProcedureDone RPCs could have been 
significantly delayed/deadlocked in the queue behind Move RPCs from the 
RegionMover, causing stale servers to reappear in the RegionStates.serverMap 
even (long) after their ServerCrashProcedure was complete. Ultimately, this is 
the same type of cause described by HBASE-23564.

However, after investigating the issue, 3.0 branch has deviated from 2.x branch 
in a way that appears unintentional, and lost the change from HBASE-23564. The 
balancer should provide its baseline of available servers from 
{{onlineServers}} instead of {{{}serverMap.keySet(){}}}.


> (branch-2) RegionStates may has some expired serverinfo and make regions do 
> not balance.
> ----------------------------------------------------------------------------------------
>
>                 Key: HBASE-30445
>                 URL: https://issues.apache.org/jira/browse/HBASE-30445
>             Project: HBase
>          Issue Type: Bug
>          Components: Balancer
>    Affects Versions: 2.6.0, 2.5.4, 2.7.0
>            Reporter: David Manning
>            Assignee: David Manning
>            Priority: Minor
>
> It is possible for the balancer to get into a state where it indefinitely 
> seeks balance but cannot obtain it. LoadBalancer logs will show
> {{balancer.StochasticLoadBalancer - Running balancer because cluster has 
> sloppy server(s).}}
> This happens indefinitely because the {{BalancerClusterState}} construction 
> adds a key twice for the same server. From here, the balancer thinks there is 
> always a server that has 0 regions, which is "sloppy", and forces a balancer 
> run even if things are otherwise balanced. The logs may look like below 
> (notice server 23 is listed twice):
> {{balancer.BalancerClusterState - server 24 is on rack 0}}
> {{balancer.BalancerClusterState - server 23 is on rack 0}}
> {{balancer.BalancerClusterState - server 23 is on rack 0}}
> {{balancer.BalancerClusterState - server 22 is on rack 0}}
> {{balancer.BalancerClusterState - server 21 is on rack 0}}
> The workaround is a restart of hmaster to clear this buggy state.
> The likely cause in this deployment, though not verified, was missing 
> HBASE-29323. As a result, ReportProcedureDone RPCs could have been 
> significantly delayed/deadlocked in the queue behind Move RPCs from the 
> RegionMover, causing stale servers to reappear in the RegionStates.serverMap 
> even (long) after their ServerCrashProcedure was complete. Ultimately, this 
> is the same type of cause described by HBASE-23564.
> However, after investigating the issue, 3.0 branch has deviated from 2.x 
> branch in a way that appears unintentional. The balancer should provide its 
> baseline of available servers from {{onlineServers}} instead of 
> {{{}serverMap.keySet(){}}}.
> 3.0 branch history:
>  # HBASE-22523 adds as {{{}serverMap.keySet(){}}}: 
> [https://github.com/apache/hbase/commit/04e5bf96d84ff2f111b1e879c8daa5360b621922#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R552-R556]
>  
>  # HBASE-23564 changes to {{{}onlineServers{}}}: 
> [https://github.com/apache/hbase/commit/ab40b9648b20cef6cb041789e4fbb11e4d36e3c2#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R562]
>  
> 2.x branch history:
>  # HBASE-22523 adds as {{{}serverMap.keySet(){}}}: 
> [https://github.com/apache/hbase/commit/14041b39f2da17de7a654a1733f3fa136fe13f63#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R552-R556]
>  # HBASE-23564 changes to {{{}onlineServers{}}}: 
> [https://github.com/apache/hbase/commit/7a0e4d814059e6a26abc834770e5f66231e0400e#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R563]
>  
>  # HBASE-23102 changes back to {{{}serverMap.keySet(){}}}: 
> [https://github.com/apache/hbase/commit/3b6d8c3394a22c71b9836c4000dbd582af248aa3#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R585]
>  



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to