[
https://issues.apache.org/jira/browse/IGNITE-21918?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Iurii Gerzhedovich updated IGNITE-21918:
----------------------------------------
Labels: MakeTeamcityGreenAgain ignite-3 (was: ignite-3)
> ItPlacementDriverReplicaSideTest#testNotificationToPlacementDriverAboutChangeLeader
> is flaky
> --------------------------------------------------------------------------------------------
>
> Key: IGNITE-21918
> URL: https://issues.apache.org/jira/browse/IGNITE-21918
> Project: Ignite
> Issue Type: Bug
> Reporter: Alexander Lapin
> Priority: Major
> Labels: MakeTeamcityGreenAgain, ignite-3
>
> {code:java}
> java.lang.AssertionError: There are replicas alive [replicas=[group_1]]
> at
> org.apache.ignite.internal.replicator.ReplicaManager.stop(ReplicaManager.java:658)
> at
> org.apache.ignite.internal.replicator.ItPlacementDriverReplicaSideTest.lambda$beforeTest$3(ItPlacementDriverReplicaSideTest.java:200)
> {code}
> Flaky rate is relatively high, nearest failures:
> #
> [02/04/24|https://ci.ignite.apache.org/buildConfiguration/ApacheIgnite3xGradle_Test_IntegrationTests_ModuleReplicator/7987165?expandBuildDeploymentsSection=false&hideTestsFromDependencies=false&expandBuildTestsSection=true&hideProblemsFromDependencies=false&expandBuildProblemsSection=true&expandBuildChangesSection=true]
> #
> [29/03/24|https://ci.ignite.apache.org/buildConfiguration/ApacheIgnite3xGradle_Test_IntegrationTests_ModuleReplicator/7996550?expandBuildDeploymentsSection=false&hideTestsFromDependencies=false&expandBuildTestsSection=true&hideProblemsFromDependencies=false&expandBuildProblemsSection=true]
> #
> [26/03/24|https://ci.ignite.apache.org/buildConfiguration/ApacheIgnite3xGradle_Test_IntegrationTests_ModuleReplicator/7973057?expandBuildDeploymentsSection=false&hideTestsFromDependencies=false&hideProblemsFromDependencies=false&expandBuildProblemsSection=true&expandBuildChangesSection=true&expandBuildTestsSection=true]
> Exception may differ a bit, but semantically it's the same
> {code:java}
> org.opentest4j.AssertionFailedError: expected: <true> but was: <false>
> at
> app//org.junit.jupiter.api.AssertionFailureBuilder.build(AssertionFailureBuilder.java:151)
> at
> app//org.junit.jupiter.api.AssertionFailureBuilder.buildAndThrow(AssertionFailureBuilder.java:132)
> at app//org.junit.jupiter.api.AssertTrue.failNotTrue(AssertTrue.java:63)
> at app//org.junit.jupiter.api.AssertTrue.assertTrue(AssertTrue.java:36)
> at app//org.junit.jupiter.api.AssertTrue.assertTrue(AssertTrue.java:31)
> at app//org.junit.jupiter.api.Assertions.assertTrue(Assertions.java:183)
> at
> app//org.apache.ignite.internal.replicator.ItPlacementDriverReplicaSideTest.testNotificationToPlacementDriverAboutChangeLeader(ItPlacementDriverReplicaSideTest.java:301)
> at [email protected]/java.lang.reflect.Method.invoke(Method.java:566)
> at [email protected]/java.util.ArrayList.forEach(ArrayList.java:1541)
> at [email protected]/java.util.ArrayList.forEach(ArrayList.java:1541)
> Suppressed: java.lang.AssertionError: There are replicas alive
> [replicas=[group_1]] {code}
> Reproduced locally 1 out of 20.
> h3. Upd#1
> Seems that it's a test issue. It's expected that three replicas will be
> started: one per each test node
> {code:java}
> [2024-04-02T11:00:46,956][INFO ][Test
> worker][ItPlacementDriverReplicaSideTest] Replication group is based on
> [ipdrst_tntpdacl_1238, ipdrst_tntpdacl_1236, ipdrst_tntpdacl_1234] {code}
> However there are only two log entries with "Replica is about to start"
> {code:java}
> [2024-04-02T11:00:48,404][INFO
> ][%ipdrst_tntpdacl_1238%JRaft-Request-Processor-8][ReplicaManager] Replica is
> about to start [replicationGroupId=group_1].
> [2024-04-02T11:00:49,071][INFO
> ][%ipdrst_tntpdacl_1234%Raft-Group-Client-2][ReplicaManager] Replica is about
> to start [replicationGroupId=group_1]. {code}
> Basically there's no "Replica is about to start" on 1236 node.
>
> Seems that there's no linearization between replica start and rest of test
> flow including replicaManager stop in @AfterEach.
> {code:java}
> private CompletableFuture<TopologyAwareRaftGroupService>
> createReplicationGroup(
> ReplicationGroupId groupId,
> Set<String> nodes
> ) throws Exception {
> var res = new CompletableFuture<TopologyAwareRaftGroupService>();
> for (String nodeName : nodes) {
> var replicaManager = replicaManagers.get(nodeName);
> var raftManager = raftManagers.get(nodeName);
> assertNotNull(replicaManager);
> assertNotNull(raftManager);
> var peer = new Peer(nodeName);
> var rftNodeId = new RaftNodeId(groupId, peer);
> CompletableFuture<TopologyAwareRaftGroupService> raftClientFut =
> raftManager.startRaftGroupNode(
> rftNodeId,
> fromConsistentIds(nodes),
> new TestRaftGroupListener(),
> RaftGroupEventsListener.noopLsnr,
> RaftGroupOptions.defaults(),
> raftClientFactory.get(nodeName)
> );
> raftClientFut.thenAccept(raftClient -> {
> try {
> if (!res.isDone()) {
> res.complete(raftClient);
> }
> replicaManager.startReplica(
> groupId,
> (request, senderId) -> {
> log.info("Handle request [type={}]",
> request.getClass().getSimpleName());
> return
> raftClient.run(REPLICA_MESSAGES_FACTORY.safeTimeSyncCommand().build())
> .thenApply(ignored -> new
> ReplicaResult(null, null));
> },
> raftClient,
> new PendingComparableValuesTracker<>(Long.MAX_VALUE));
> } catch (NodeStoppingException e) {
> fail("Can not start replica [groupId=" + groupId + ']');
> }
> });
> }
> return res;
> } {code}
> Basically createReplicationGroup completes the future with raftClient
> ignoring the fact that replicaManager start is itself asynchronous.
>
--
This message was sent by Atlassian Jira
(v8.20.10#820010)