[
https://issues.apache.org/jira/browse/HBASE-30445?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
David Manning updated HBASE-30445:
----------------------------------
Description:
It is possible for the balancer to get into a state where it indefinitely seeks
balance but cannot obtain it. LoadBalancer logs will show
{{balancer.StochasticLoadBalancer - Running balancer because cluster has sloppy
server(s).}}
This happens indefinitely because the {{BalancerClusterState}} construction
adds a key twice for the same server. From here, the balancer thinks there is
always a server that has 0 regions, which is "sloppy", and forces a balancer
run even if things are otherwise balanced. The logs may look like below (notice
server 23 is listed twice):
{{balancer.BalancerClusterState - server 24 is on rack 0}}
{{balancer.BalancerClusterState - server 23 is on rack 0}}
{{balancer.BalancerClusterState - server 23 is on rack 0}}
{{balancer.BalancerClusterState - server 22 is on rack 0}}
{{balancer.BalancerClusterState - server 21 is on rack 0}}
The workaround is a restart of hmaster to clear this buggy state.
The likely cause in this deployment, though not verified, was missing
HBASE-29323. As a result, ReportProcedureDone RPCs could have been
significantly delayed/deadlocked in the queue behind Move RPCs from the
RegionMover, causing stale servers to reappear in the RegionStates.serverMap
even (long) after their ServerCrashProcedure was complete. Ultimately, this is
the same type of cause described by HBASE-23564.
However, after investigating the issue, 3.0 branch has deviated from 2.x branch
in a way that appears unintentional. The balancer should provide its baseline
of available servers from {{onlineServers}} instead of
{{{}serverMap.keySet(){}}}.
3.0 branch history:
# HBASE-22523 adds as {{{}serverMap.keySet(){}}}:
[https://github.com/apache/hbase/commit/04e5bf96d84ff2f111b1e879c8daa5360b621922#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R552-R556]
# HBASE-23564 changes to {{{}onlineServers{}}}:
[https://github.com/apache/hbase/commit/ab40b9648b20cef6cb041789e4fbb11e4d36e3c2#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R562]
2.x branch history:
# HBASE-22523 adds as {{{}serverMap.keySet(){}}}:
[https://github.com/apache/hbase/commit/14041b39f2da17de7a654a1733f3fa136fe13f63#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R552-R556]
# HBASE-23564 changes to {{{}onlineServers{}}}:
[https://github.com/apache/hbase/commit/7a0e4d814059e6a26abc834770e5f66231e0400e#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R563]
# HBASE-23102 changes back to {{{}serverMap.keySet(){}}}:
[https://github.com/apache/hbase/commit/3b6d8c3394a22c71b9836c4000dbd582af248aa3#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R585]
was:
It is possible for the balancer to get into a state where it indefinitely seeks
balance but cannot obtain it. LoadBalancer logs will show
{{balancer.StochasticLoadBalancer - Running balancer because cluster has sloppy
server(s).}}
This happens indefinitely because the {{BalancerClusterState}} construction
adds a key twice for the same server. From here, the balancer thinks there is
always a server that has 0 regions, which is "sloppy", and forces a balancer
run even if things are otherwise balanced. The logs may look like below (notice
server 23 is listed twice):
{{balancer.BalancerClusterState - server 24 is on rack 0}}
{{balancer.BalancerClusterState - server 23 is on rack 0}}
{{balancer.BalancerClusterState - server 23 is on rack 0}}
{{balancer.BalancerClusterState - server 22 is on rack 0}}
{{balancer.BalancerClusterState - server 21 is on rack 0}}
The workaround is a restart of hmaster to clear this buggy state.
The likely cause in this deployment, though not verified, was missing
HBASE-29323. As a result, ReportProcedureDone RPCs could have been
significantly delayed/deadlocked in the queue behind Move RPCs from the
RegionMover, causing stale servers to reappear in the RegionStates.serverMap
even (long) after their ServerCrashProcedure was complete. Ultimately, this is
the same type of cause described by HBASE-23564.
However, after investigating the issue, 3.0 branch has deviated from 2.x branch
in a way that appears unintentional, and lost the change from HBASE-23564. The
balancer should provide its baseline of available servers from
{{onlineServers}} instead of {{{}serverMap.keySet(){}}}.
> (branch-2) RegionStates may has some expired serverinfo and make regions do
> not balance.
> ----------------------------------------------------------------------------------------
>
> Key: HBASE-30445
> URL: https://issues.apache.org/jira/browse/HBASE-30445
> Project: HBase
> Issue Type: Bug
> Components: Balancer
> Affects Versions: 2.6.0, 2.5.4, 2.7.0
> Reporter: David Manning
> Assignee: David Manning
> Priority: Minor
>
> It is possible for the balancer to get into a state where it indefinitely
> seeks balance but cannot obtain it. LoadBalancer logs will show
> {{balancer.StochasticLoadBalancer - Running balancer because cluster has
> sloppy server(s).}}
> This happens indefinitely because the {{BalancerClusterState}} construction
> adds a key twice for the same server. From here, the balancer thinks there is
> always a server that has 0 regions, which is "sloppy", and forces a balancer
> run even if things are otherwise balanced. The logs may look like below
> (notice server 23 is listed twice):
> {{balancer.BalancerClusterState - server 24 is on rack 0}}
> {{balancer.BalancerClusterState - server 23 is on rack 0}}
> {{balancer.BalancerClusterState - server 23 is on rack 0}}
> {{balancer.BalancerClusterState - server 22 is on rack 0}}
> {{balancer.BalancerClusterState - server 21 is on rack 0}}
> The workaround is a restart of hmaster to clear this buggy state.
> The likely cause in this deployment, though not verified, was missing
> HBASE-29323. As a result, ReportProcedureDone RPCs could have been
> significantly delayed/deadlocked in the queue behind Move RPCs from the
> RegionMover, causing stale servers to reappear in the RegionStates.serverMap
> even (long) after their ServerCrashProcedure was complete. Ultimately, this
> is the same type of cause described by HBASE-23564.
> However, after investigating the issue, 3.0 branch has deviated from 2.x
> branch in a way that appears unintentional. The balancer should provide its
> baseline of available servers from {{onlineServers}} instead of
> {{{}serverMap.keySet(){}}}.
> 3.0 branch history:
> # HBASE-22523 adds as {{{}serverMap.keySet(){}}}:
> [https://github.com/apache/hbase/commit/04e5bf96d84ff2f111b1e879c8daa5360b621922#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R552-R556]
>
> # HBASE-23564 changes to {{{}onlineServers{}}}:
> [https://github.com/apache/hbase/commit/ab40b9648b20cef6cb041789e4fbb11e4d36e3c2#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R562]
>
> 2.x branch history:
> # HBASE-22523 adds as {{{}serverMap.keySet(){}}}:
> [https://github.com/apache/hbase/commit/14041b39f2da17de7a654a1733f3fa136fe13f63#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R552-R556]
> # HBASE-23564 changes to {{{}onlineServers{}}}:
> [https://github.com/apache/hbase/commit/7a0e4d814059e6a26abc834770e5f66231e0400e#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R563]
>
> # HBASE-23102 changes back to {{{}serverMap.keySet(){}}}:
> [https://github.com/apache/hbase/commit/3b6d8c3394a22c71b9836c4000dbd582af248aa3#diff-96d9ec3583f3683e032c43a8f2561e06888bf1fe5eb4e818b3fae47b379ef196R585]
>
--
This message was sent by Atlassian Jira
(v8.20.10#820010)