Maksim Davydov created IGNITE-29092:
---------------------------------------
Summary: Fix flaky
IgniteClusterSnapshotStreamerTest.testMetaWarningRestoredByOnlyOneNode
Key: IGNITE-29092
URL: https://issues.apache.org/jira/browse/IGNITE-29092
Project: Ignite
Issue Type: Test
Reporter: Maksim Davydov
Assignee: Maksim Davydov
IgniteClusterSnapshotStreamerTest.testMetaWarningRestoredByOnlyOneNode fails
from time to time in the Snapshots 3 suite with:
ClusterTopologyCheckedException: Snapshot validation stopped. Required node has
left the cluster [nodeId=[...0001]]
It fails with encryption on and off, for example:
-
https://ci2.ignite.apache.org/buildConfiguration/IgniteTests24Java8_Snapshots3/9375047
-
https://ci2.ignite.apache.org/buildConfiguration/IgniteTests24Java8_Snapshots3/9373027
The test stops servers 0 and 1 and then checks the snapshot from the client.
SnapshotCheckProcess#start takes the required nodes from the local discovery
cache of the node that starts the check, here the client
(aliveBaselineNodes()). stopGrid() waits until each node's discovery and
exchange versions match, but not until other nodes have processed the leave. If
the client hasn't processed NODE_LEFT for node 1 yet, the check lists node 1 as
required. Node 1 never answers, and the check fails. In both builds the failure
names the server stopped last.
The race reproduces every time if the client's processing of that NODE_LEFT is
delayed by 2 seconds: 12 of 12 executions failed with the same error.
The fix: before the check, wait until every node, the client included, sees the
new topology (waitForTopology(3)). With the same 2-second delay, 11 of 11
executions pass.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)