Gabriella Lotz created KUDU-3808:
------------------------------------
Summary: Auto leader rebalancer should send LeaderStepDown RPCs
asynchronously
Key: KUDU-3808
URL: https://issues.apache.org/jira/browse/KUDU-3808
Project: Kudu
Issue Type: Improvement
Reporter: Gabriella Lotz
{{AutoLeaderRebalancerTask}} sends its {{LeaderStepDown}} RPCs one at a time
and blocks on each reply. Both passes do it this way:
{{RunLeaderRebalanceForTable()}} and {{{}RunGlobalLeaderRebalance(){}}}.
Each call gets {{-auto_leader_rebalancing_rpc_timeout_seconds}} (10s by
default), and a round sends at most {{leader_rebalancing_max_moves_per_round}}
(10 by default), so a round whose targets are slow to answer can spend around
100s just waiting. That's small next to the default
{{-auto_leader_rebalancing_interval_seconds}} of 3600, but it grows with the
move limit, so raising that flag makes rounds proportionally slower.
Proposal: dispatch the step-downs asynchronously, and wait for the outstanding
calls before the round finishes. {{AsyncLeaderStepDownForReplacementTask}} in
catalog_manager.cc already does this on top of {{RetryingTSRpcTask}} and
{{{}ConsensusServiceProxy::LeaderStepDownAsync(){}}}, so there's a pattern to
follow.
A few things to keep in mind:
- The counters added in KUDU-3791
({{{}auto_leader_rebalancer_moves_scheduled{}}}, {{{}_moves_completed{}}},
{{{}_moves_failed{}}}) are bumped inline today. They'd move into the response
callbacks, and the invariant {{scheduled = completed + failed}} needs to keep
holding when moves complete concurrently.
- Those same counters give us a way to check the async version isn't losing
moves. {{LeaderRebalancerTest.RebalancerMetrics}} already asserts the invariant.
- A round shouldn't return while step-downs are still in flight, otherwise the
next round replans against a cluster that's still moving.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)