Maksim Davydov created IGNITE-29088:
---------------------------------------
Summary: Coordinator keeps a lost partition MOVING after its new
primary owns it (IGNORE loss policy)
Key: IGNITE-29088
URL: https://issues.apache.org/jira/browse/IGNITE-29088
Project: Ignite
Issue Type: Bug
Reporter: Maksim Davydov
Assignee: Maksim Davydov
*Problem*
In an in-memory cluster with the IGNORE partition loss policy (baseline
auto-adjust on, timeout 0), a node leaves and takes the only copy of some
partitions with it (backups = 0). Each lost partition is recreated empty on its
new primary and owned there. Sometimes the coordinator never learns that: its
partition map keeps the partition MOVING until the next exchange.
*Impact*
Until the next topology change, which in a stable cluster can be hours away,
the coordinator and every node that takes its full map see no owner of the
partition:
- SQL queries from those nodes fail after the retry timeout:
"Failed to map SQL query to topology during timeout: 30000ms". With a MOVING
partition in the map, ReducePartitionMapper maps partitions by their owners and
finds none. On the new primary, whose own map is right, the same query succeeds.
- awaitPartitionMapExchange() in tests times out. Key operations and scan
queries still work.
This is not what IGNORE is meant to do: the policy resets a lost partition
silently, no
partition is LOST, and there is nothing to reset by hand.
*Cause*
With the exchange merge protocol, the new primary creates the lost partition as
MOVING when it gets the coordinator's full message, and owns it a moment later
in detectLostPartitions. The coordinator, in its own detectLostPartitions,
already marks the partition OWNING for that node.
If the new primary sends its partition map between these two steps, the map
carries MOVING with a newer update sequence, and the coordinator takes it. A
typical sender is the scheduled resend (scheduleResendPartitions(), 1.5 s after
an eviction or a map change). Nothing sends the node's map again after it owns
the partition: the result of detectLostPartitions is dropped in
GridDhtPartitionsExchangeFuture#detectLostPartitions.
Also, inside GridDhtPartitionTopologyImpl#detectLostPartitions the result
reflects only the last lost partition.
*How it shows*
IgniteTopologyValidatorGridSplitCacheTest (32 nodes, 50 caches without backups)
times out in awaitPartitionMapExchange() after stopping its configless node in
9 of 15 local runs. Small clusters rarely send a map inside that window, so
IgniteCachePartitionLossPolicySelfTest (3 nodes) doesn't hit it.
*Fix*
When detectLostPartitions changes a local partition, a non-coordinator node
sends its single map again. A new test, CachePartitionLossIgnorePolicyMapTest,
makes the new primary send its map inside the window every time and checks that
the coordinator's map matches the nodes.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)