Chu Cheng Li created HDDS-16623:
-----------------------------------

             Summary: Datanode can be missing from NetworkTopology after it 
recovers from DEAD
                 Key: HDDS-16623
                 URL: https://issues.apache.org/jira/browse/HDDS-16623
             Project: Apache Ozone
          Issue Type: Improvement
            Reporter: Chu Cheng Li


SCM removes a datanode from its NetworkTopology when the datanode becomes DEAD, 
and adds it back when the datanode starts heartbeating again and moves to 
HEALTHY_READONLY. Since HDDS-5437 this has been done by two event handlers: 
DeadNodeHandler removes the node and HealthyReadOnlyNodeHandler adds it back. 
The handlers run on separate EventQueue threads, and each one acts on a node 
state it read a little earlier.

HDDS-14834 made DeadNodeHandler check the node state again just before removing 
the node. The check and the removal are still separate steps, though, so this 
can happen:

1. DeadNodeHandler checks the node and sees it is still DEAD.
2. The datanode heartbeats again. SCM moves it to HEALTHY_READONLY and fires 
HEALTHY_READONLY_NODE.
3. HealthyReadOnlyNodeHandler adds the node to the topology. Nothing changes, 
because the node is still there.
4. DeadNodeHandler removes the node.

The datanode is now HEALTHY_READONLY, and soon HEALTHY, but it is missing from 
the topology. Placement policies that pick nodes from the topology won't choose 
it until it next moves to HEALTHY_READONLY.

*Proposed fix*: update the topology in NodeStateManager at the same time as the 
health state changes. Remove the node when it becomes DEAD, and add it back 
when it recovers. Both changes happen inside the synchronized health check, so 
the topology always changes in the same order as the node's health. The event 
handlers no
longer need to touch it.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to