peterxcli opened a new pull request, #11343:
URL: https://github.com/apache/ozone/pull/11343
## What changes were proposed in this pull request?
SCM removes a datanode from its NetworkTopology when the datanode becomes
DEAD, and adds it back when the
datanode starts heartbeating again and moves to HEALTHY_READONLY. Until now
two event handlers did this:
DeadNodeHandler removed the node and HealthyReadOnlyNodeHandler added it
back. They run on separate
EventQueue threads and act on a node state they read a little earlier, so
their updates can land in the
wrong order. HDDS-14834 narrowed the window with an extra state check in
DeadNodeHandler, but this can
still happen:
1. DeadNodeHandler checks the node and sees it is still DEAD.
2. The datanode heartbeats again. SCM moves it to HEALTHY_READONLY and fires
HEALTHY_READONLY_NODE.
3. HealthyReadOnlyNodeHandler adds the node to the topology. Nothing
changes, because the node is still there.
4. DeadNodeHandler removes the node.
The datanode is then healthy but missing from the topology, so placement
policies that pick nodes from the
topology won't choose it.
This PR moves the topology update into NodeStateManager, next to the health
state change itself. The node is
removed when it becomes DEAD and added back when it recovers. Both
transitions happen inside the synchronized
health check, so the topology always changes in the same order as the node's
health. The handlers no longer
touch the topology. That also removes the opposite race, where a late add
puts back a node that has died
again, and the parent-pointer checks in both handlers that could fail during
the race. A failure while
updating the topology is logged rather than thrown, because an exception
there would stop SCM from
scheduling further health checks.
Test changes:
- TestNodeStateManager: new tests for the topology updates when a node
becomes DEAD and recovers, and for a
failing topology update not breaking the health check.
- TestDeadNodeHandler: removes the HDDS-14834 tests for handler behaviour
that no longer exists. The
remaining one now checks that DeadNodeHandler skips a node that is no
longer DEAD.
- TestQueryNode: new end-to-end test that stops a datanode, waits for it to
leave SCM's topology, then
restarts it and checks it is back.
## What is the link to the Apache JIRA
https://issues.apache.org/jira/browse/HDDS-16623
## How was this patch tested?
- All tests in the SCM `org.apache.hadoop.hdds.scm.node` package pass, except
`TestSCMNodeManager#testScmClusterIsInExpectedState2`, which is
timing-sensitive and also fails
intermittently on master (2 of 3 local runs on unmodified master).
- `TestQueryNode` in `ozone-integration-test`, including the new end-to-end
test.
- checkstyle and PMD.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]