zstan commented on code in PR #13130:
URL: https://github.com/apache/ignite/pull/13130#discussion_r3631950517


##########
docs/_docs/perf-and-troubleshooting/general-perf-tips.adoc:
##########
@@ -47,3 +47,458 @@ queries with JOINs at massive scale and expect significant 
performance benefits.
 
 * Adjust link:data-rebalancing[data rebalancing settings] to ensure that 
rebalancing completes faster when your cluster topology changes.
 
+== How to assess cluster health
+
+Cluster health is a complex thing. Apache Ignite is capable of demonstrating 
great performance across different scenarios with varying loads. Therefore, in 
general terms, a healthy cluster is one whose behavior aligns with your 
expectations. However, there are some universal aspects that apply to all 
deployments and warrant attention.
+
+It is important to understand that a healthy cluster may undergo planned 
topology changes or temporary load spikes.
+
+The key properties are:
+
+* The cluster is in the intended link:monitoring-metrics/cluster-states[state] 
and serves only the operations allowed by that state.
+* Baseline topology, when it is used or managed manually, matches the expected 
data-bearing server nodes.
+* Data remains consistent: 
link:tools/control-script#verifying-partition-checksums[`idle_verify`] reports 
no partition conflicts when the cluster is idle.
+* Expected nodes are present, and no unexpected repeated node `JOIN`, `LEFT`, 
or `FAIL` events or node segmentation are reported.
+* link:configuring-caches/partition-loss-policy[Lost partitions] are absent.
+* The number and age of 
link:tools/control-script#transaction-management[long-running transactions] do 
not keep increasing; active 
link:data-modeling/data-partitioning#partition-map-exchange[Partition Map 
Exchanges (PMEs)] complete; 
link:monitoring-metrics/new-metrics-system#monitoring-rebalancing[rebalancing] 
finishes; 
link:monitoring-metrics/new-metrics-system#monitoring-checkpointing-operations[checkpoint-related
 metrics] and link:monitoring-metrics/new-metrics#thread-pools[executor queue 
sizes] return to their usual ranges after the 
link:monitoring-metrics/new-metrics-system#monitoring-topology[topology] has 
stabilized and the application workload has returned to its expected level.
+There is no single command or metric that proves cluster health for every 
deployment.
+
+Use several signals together.
+A simple client connection or SQL liveness check can prove only that a 
particular client or query path is reachable; it does not check user 
partitions, backup consistency, baseline membership, or all server nodes.
+
+=== Check the Intended State and Node Membership
+
+Start with the link:monitoring-metrics/cluster-states[cluster state].
+
+Run:
+
+[source,shell]
+----
+control.(sh|bat) --state
+----
+
+Relevant output:
+
+[source,text]
+----
+Command [STATE] started
+Arguments: --state
+--------------------------------------------------------------------------------
+Cluster state: ACTIVE
+Command [STATE] finished with code: 0
+----
+
+`ACTIVE` is expected for normal read-write operation.
+`ACTIVE_READ_ONLY` is normal when read-only operation was intentionally 
enabled.
+`INACTIVE` is acceptable only when it matches the current operation, for 
example planned maintenance; an inactive cluster does not serve the data 
workload.
+The criterion is whether the actual state matches the state that was 
intentionally set for the deployment.
+
+Check baseline topology when the cluster uses persistence, when baseline 
autoadjustment is disabled, or when you intentionally manage the set of 
data-bearing nodes.
+In pure in-memory clusters with the default immediate autoadjustment, baseline 
topology normally follows the current server topology automatically.
+If autoadjustment is disabled, the baseline changes only after an operator 
changes it.
+If autoadjustment is configured with a non-zero timeout, the baseline is 
updated only after the topology remains unchanged for that timeout.
+In both cases, run `control.(sh|bat) --baseline` and compare `Baseline nodes` 
and `Other nodes` with the expected set of server nodes.
+
+Run:
+
+[source,shell]
+----
+control.(sh|bat) --baseline
+----
+
+Relevant output for a cluster where all baseline nodes are online:
+
+[source,text]
+----
+Cluster state: ACTIVE
+Current topology version: 3
+Baseline auto adjustment disabled: softTimeout=300000
+
+Current topology version: 3 (Coordinator: ConsistentId=node-1, Order=1)
+
+Baseline nodes:
+    ConsistentId=node-1, State=ONLINE, Order=1
+    ConsistentId=node-2, State=ONLINE, Order=2
+    ConsistentId=node-3, State=ONLINE, Order=3
+--------------------------------------------------------------------------------
+Number of baseline nodes: 3
+
+Other nodes not found.
+----
+
+Example: one baseline node is offline:
+
+[source,text]
+----
+Baseline nodes:
+    ConsistentId=node-1, State=ONLINE, Order=1
+    ConsistentId=node-2, State=OFFLINE, Order=2
+    ConsistentId=node-3, State=ONLINE, Order=3
+--------------------------------------------------------------------------------
+Number of baseline nodes: 3
+----
+
+If a baseline node is `OFFLINE`, an expected data-bearing server is missing.
+If other primary or backup copies are available, its absence does not cause 
partition loss.
+To check for partition loss, query the partition states as described in 
<<confirm-that-rebalancing-converges,Confirm That Rebalancing Converges>>.
+
+If an online server node has joined the cluster but is not in the baseline, 
the command shows it under `Other nodes`:
+
+[source,text]
+----
+Other nodes:
+    ConsistentId=node-4, Order=4
+Number of other nodes: 1
+----
+
+The baseline contains server nodes that are intended to store a data.
+Client nodes are not a part of the baseline.
+An online server node in `Other nodes` is not always an error: the node may 
have been prepared intentionally but not yet introduced into the data topology.
+If the node is expected to store data, first check the 
link:clustering/baseline-topology#baseline-topology-autoadjustment[baseline 
auto-adjustment policy] and the current maintenance or scale-out procedure, 
then use the documented baseline change procedure.
+Changing the baseline can start link:data-rebalancing[rebalancing]: partitions 
are redistributed according to the new affinity assignment.
+Plan for the additional network, CPU, and storage load, especially in clusters 
with persistence.
+
+Use topology changes to distinguish planned activity from instability.
+A server `JOIN`, `LEFT`, or `FAIL` event changes cluster membership and 
triggers link:data-modeling/data-partitioning#partition-map-exchange[Partition 
Map Exchanges (PME)].
+A dynamic cache start or stop can also trigger PME even when no server node 
joins, leaves, or fails.
+Therefore, a topology version or 
link:monitoring-metrics/new-metrics#partition-map-exchange[PME metric] PME 
metric change is useful only when interpreted together with maintenance actions 
and node logs.
+For example, these metrics should be used for historical distributions; there 
is no universal duration threshold.
+
+[source,sql]
+----
+SELECT NAME, VALUE
+FROM SYS.METRICS
+WHERE NAME IN (
+    'pme.Duration',
+    'pme.CacheOperationsBlockedDuration'
+)
+ORDER BY NAME;
+----
+
+Run `control.(sh|bat) --baseline` repeatedly or monitor topology metrics to 
confirm that membership is stable when no maintenance is in progress.
+In logs, look for repeated node join, left, fail, segmentation, and 
exchange-worker messages.
+Investigate unexpected repeated `JOIN`, `LEFT`, or `FAIL` events, node 
segmentation, network failures, or a PME that does not finish.
+For PME metrics and transaction checks, see 
<<check-transactions-and-sql-queries,Check Transactions and SQL Queries>>.
+
+=== Verify Partition Consistency
+
+When the cluster is expected to be idle, run:
+
+[source,shell]
+----
+control.(sh|bat) --cache idle_verify
+----
+
+Successful result:
+
+[source,text]
+----
+The check procedure has finished, no conflicts have been found.
+----
+
+The beginning of a conflict result uses this format:
+
+[source,text]
+----
+The check procedure has failed, conflict partitions has been found: 
[counterConflicts=1, hashConflicts=0]
+Update counter conflicts:
+Conflict partition: PartitionKey [grpId=1544803905, grpName=default, partId=5]
+----
+
+The command compares partition update counters and partition hashes between 
primary and backup copies.
+Run it only when data updates are stopped.
+If updates are active, the command can report false conflicts because copies 
are changing while hashes are being calculated.
+Partitions in `MOVING` or `LOST` state may be skipped, so the result can be 
incomplete.
+A successful `idle_verify` result is an important confirmation of consistency, 
but it still does not prove overall cluster health check success.
+
+[#confirm-that-rebalancing-converges]
+=== Confirm That Rebalancing Converges
+
+Immediately after an intended topology or cache event, `MOVING` and `RENTING` 
counts can be non-zero.
+After the topology becomes stable, repeat the query and confirm that both 
counts decrease and eventually disappear.
+A `LOST` count greater than zero is not a normal transient rebalance state.
+
+[source,sql]
+----
+SELECT STATE, COUNT(*) AS PARTITION_COUNT
+FROM SYS.PARTITION_STATES
+WHERE STATE IN ('MOVING', 'RENTING', 'LOST')
+GROUP BY STATE
+ORDER BY STATE;
+----
+
+It is also possible to track some node-local cache-group metrics, such as:
+
+* LocalNodeMovingPartitionsCount;
+* LocalNodeRentingPartitionsCount;
+* LocalNodeRentingEntriesCount.
+
+Use the 
link:monitoring-metrics/system-views#partition_states[PARTITION_STATES] system 
view to check partition states:
+
+* `OWNING`: the node is the current primary or backup owner.
+* `MOVING`: a partition copy is being loaded on the node during rebalance.
+* `RENTING`: an old copy is being removed after ownership changes.
+* `EVICTED`: the partition is absent on a node that is no longer an owner; 
this is not an error by itself.
+* `LOST`: the partition is unavailable and must not be used; investigate 
immediately.
+
+[source,sql]
+----
+SELECT CACHE_GROUP_ID, PARTITION_ID, NODE_ID, STATE, IS_PRIMARY
+FROM SYS.PARTITION_STATES
+WHERE STATE IN ('MOVING', 'RENTING', 'LOST')
+ORDER BY STATE, CACHE_GROUP_ID, PARTITION_ID, NODE_ID;
+----
+
+In steady state, this query usually should not return `MOVING`, `RENTING`, or 
`LOST` rows.

Review Comment:
   ```suggestion
   In stable cluster state, this query usually should not return `MOVING`, 
`RENTING`, or `LOST` rows.
   ```



##########
docs/_docs/perf-and-troubleshooting/general-perf-tips.adoc:
##########
@@ -47,3 +47,458 @@ queries with JOINs at massive scale and expect significant 
performance benefits.
 
 * Adjust link:data-rebalancing[data rebalancing settings] to ensure that 
rebalancing completes faster when your cluster topology changes.
 
+== How to assess cluster health
+
+Cluster health is a complex thing. Apache Ignite is capable of demonstrating 
great performance across different scenarios with varying loads. Therefore, in 
general terms, a healthy cluster is one whose behavior aligns with your 
expectations. However, there are some universal aspects that apply to all 
deployments and warrant attention.
+
+It is important to understand that a healthy cluster may undergo planned 
topology changes or temporary load spikes.
+
+The key properties are:
+
+* The cluster is in the intended link:monitoring-metrics/cluster-states[state] 
and serves only the operations allowed by that state.
+* Baseline topology, when it is used or managed manually, matches the expected 
data-bearing server nodes.
+* Data remains consistent: 
link:tools/control-script#verifying-partition-checksums[`idle_verify`] reports 
no partition conflicts when the cluster is idle.
+* Expected nodes are present, and no unexpected repeated node `JOIN`, `LEFT`, 
or `FAIL` events or node segmentation are reported.
+* link:configuring-caches/partition-loss-policy[Lost partitions] are absent.
+* The number and age of 
link:tools/control-script#transaction-management[long-running transactions] do 
not keep increasing; active 
link:data-modeling/data-partitioning#partition-map-exchange[Partition Map 
Exchanges (PMEs)] complete; 
link:monitoring-metrics/new-metrics-system#monitoring-rebalancing[rebalancing] 
finishes; 
link:monitoring-metrics/new-metrics-system#monitoring-checkpointing-operations[checkpoint-related
 metrics] and link:monitoring-metrics/new-metrics#thread-pools[executor queue 
sizes] return to their usual ranges after the 
link:monitoring-metrics/new-metrics-system#monitoring-topology[topology] has 
stabilized and the application workload has returned to its expected level.

Review Comment:
   ```suggestion
   * The number and timeout of 
link:tools/control-script#transaction-management[long-running transactions] do 
not keep increasing; active 
link:data-modeling/data-partitioning#partition-map-exchange[Partition Map 
Exchanges (PMEs)] complete; 
link:monitoring-metrics/new-metrics-system#monitoring-rebalancing[rebalancing] 
finishes; 
link:monitoring-metrics/new-metrics-system#monitoring-checkpointing-operations[checkpoint-related
 metrics] and link:monitoring-metrics/new-metrics#thread-pools[executor queue 
sizes] return to their usual ranges after the 
link:monitoring-metrics/new-metrics-system#monitoring-topology[topology] has 
stabilized and the application workload has returned to its expected level.
   ```



##########
docs/_docs/perf-and-troubleshooting/general-perf-tips.adoc:
##########
@@ -47,3 +47,458 @@ queries with JOINs at massive scale and expect significant 
performance benefits.
 
 * Adjust link:data-rebalancing[data rebalancing settings] to ensure that 
rebalancing completes faster when your cluster topology changes.
 
+== How to assess cluster health
+
+Cluster health is a complex thing. Apache Ignite is capable of demonstrating 
great performance across different scenarios with varying loads. Therefore, in 
general terms, a healthy cluster is one whose behavior aligns with your 
expectations. However, there are some universal aspects that apply to all 
deployments and warrant attention.
+
+It is important to understand that a healthy cluster may undergo planned 
topology changes or temporary load spikes.
+
+The key properties are:
+
+* The cluster is in the intended link:monitoring-metrics/cluster-states[state] 
and serves only the operations allowed by that state.
+* Baseline topology, when it is used or managed manually, matches the expected 
data-bearing server nodes.
+* Data remains consistent: 
link:tools/control-script#verifying-partition-checksums[`idle_verify`] reports 
no partition conflicts when the cluster is idle.
+* Expected nodes are present, and no unexpected repeated node `JOIN`, `LEFT`, 
or `FAIL` events or node segmentation are reported.
+* link:configuring-caches/partition-loss-policy[Lost partitions] are absent.
+* The number and age of 
link:tools/control-script#transaction-management[long-running transactions] do 
not keep increasing; active 
link:data-modeling/data-partitioning#partition-map-exchange[Partition Map 
Exchanges (PMEs)] complete; 
link:monitoring-metrics/new-metrics-system#monitoring-rebalancing[rebalancing] 
finishes; 
link:monitoring-metrics/new-metrics-system#monitoring-checkpointing-operations[checkpoint-related
 metrics] and link:monitoring-metrics/new-metrics#thread-pools[executor queue 
sizes] return to their usual ranges after the 
link:monitoring-metrics/new-metrics-system#monitoring-topology[topology] has 
stabilized and the application workload has returned to its expected level.
+There is no single command or metric that proves cluster health for every 
deployment.
+
+Use several signals together.
+A simple client connection or SQL liveness check can prove only that a 
particular client or query path is reachable; it does not check user 
partitions, backup consistency, baseline membership, or all server nodes.
+
+=== Check the Intended State and Node Membership
+
+Start with the link:monitoring-metrics/cluster-states[cluster state].
+
+Run:
+
+[source,shell]
+----
+control.(sh|bat) --state
+----
+
+Relevant output:
+
+[source,text]
+----
+Command [STATE] started
+Arguments: --state
+--------------------------------------------------------------------------------
+Cluster state: ACTIVE
+Command [STATE] finished with code: 0
+----
+
+`ACTIVE` is expected for normal read-write operation.
+`ACTIVE_READ_ONLY` is normal when read-only operation was intentionally 
enabled.
+`INACTIVE` is acceptable only when it matches the current operation, for 
example planned maintenance; an inactive cluster does not serve the data 
workload.
+The criterion is whether the actual state matches the state that was 
intentionally set for the deployment.
+
+Check baseline topology when the cluster uses persistence, when baseline 
autoadjustment is disabled, or when you intentionally manage the set of 
data-bearing nodes.
+In pure in-memory clusters with the default immediate autoadjustment, baseline 
topology normally follows the current server topology automatically.
+If autoadjustment is disabled, the baseline changes only after an operator 
changes it.
+If autoadjustment is configured with a non-zero timeout, the baseline is 
updated only after the topology remains unchanged for that timeout.
+In both cases, run `control.(sh|bat) --baseline` and compare `Baseline nodes` 
and `Other nodes` with the expected set of server nodes.
+
+Run:
+
+[source,shell]
+----
+control.(sh|bat) --baseline
+----
+
+Relevant output for a cluster where all baseline nodes are online:
+
+[source,text]
+----
+Cluster state: ACTIVE
+Current topology version: 3
+Baseline auto adjustment disabled: softTimeout=300000
+
+Current topology version: 3 (Coordinator: ConsistentId=node-1, Order=1)
+
+Baseline nodes:
+    ConsistentId=node-1, State=ONLINE, Order=1
+    ConsistentId=node-2, State=ONLINE, Order=2
+    ConsistentId=node-3, State=ONLINE, Order=3
+--------------------------------------------------------------------------------
+Number of baseline nodes: 3
+
+Other nodes not found.
+----
+
+Example: one baseline node is offline:
+
+[source,text]
+----
+Baseline nodes:
+    ConsistentId=node-1, State=ONLINE, Order=1
+    ConsistentId=node-2, State=OFFLINE, Order=2
+    ConsistentId=node-3, State=ONLINE, Order=3
+--------------------------------------------------------------------------------
+Number of baseline nodes: 3
+----
+
+If a baseline node is `OFFLINE`, an expected data-bearing server is missing.
+If other primary or backup copies are available, its absence does not cause 
partition loss.
+To check for partition loss, query the partition states as described in 
<<confirm-that-rebalancing-converges,Confirm That Rebalancing Converges>>.
+
+If an online server node has joined the cluster but is not in the baseline, 
the command shows it under `Other nodes`:
+
+[source,text]
+----
+Other nodes:
+    ConsistentId=node-4, Order=4
+Number of other nodes: 1
+----
+
+The baseline contains server nodes that are intended to store a data.
+Client nodes are not a part of the baseline.
+An online server node in `Other nodes` is not always an error: the node may 
have been prepared intentionally but not yet introduced into the data topology.
+If the node is expected to store data, first check the 
link:clustering/baseline-topology#baseline-topology-autoadjustment[baseline 
auto-adjustment policy] and the current maintenance or scale-out procedure, 
then use the documented baseline change procedure.
+Changing the baseline can start link:data-rebalancing[rebalancing]: partitions 
are redistributed according to the new affinity assignment.
+Plan for the additional network, CPU, and storage load, especially in clusters 
with persistence.
+
+Use topology changes to distinguish planned activity from instability.
+A server `JOIN`, `LEFT`, or `FAIL` event changes cluster membership and 
triggers link:data-modeling/data-partitioning#partition-map-exchange[Partition 
Map Exchanges (PME)].
+A dynamic cache start or stop can also trigger PME even when no server node 
joins, leaves, or fails.
+Therefore, a topology version or 
link:monitoring-metrics/new-metrics#partition-map-exchange[PME metric] PME 
metric change is useful only when interpreted together with maintenance actions 
and node logs.
+For example, these metrics should be used for historical distributions; there 
is no universal duration threshold.
+
+[source,sql]
+----
+SELECT NAME, VALUE
+FROM SYS.METRICS
+WHERE NAME IN (
+    'pme.Duration',
+    'pme.CacheOperationsBlockedDuration'
+)
+ORDER BY NAME;
+----
+
+Run `control.(sh|bat) --baseline` repeatedly or monitor topology metrics to 
confirm that membership is stable when no maintenance is in progress.
+In logs, look for repeated node join, left, fail, segmentation, and 
exchange-worker messages.
+Investigate unexpected repeated `JOIN`, `LEFT`, or `FAIL` events, node 
segmentation, network failures, or a PME that does not finish.
+For PME metrics and transaction checks, see 
<<check-transactions-and-sql-queries,Check Transactions and SQL Queries>>.
+
+=== Verify Partition Consistency
+
+When the cluster is expected to be idle, run:
+
+[source,shell]
+----
+control.(sh|bat) --cache idle_verify
+----
+
+Successful result:
+
+[source,text]
+----
+The check procedure has finished, no conflicts have been found.
+----
+
+The beginning of a conflict result uses this format:
+
+[source,text]
+----
+The check procedure has failed, conflict partitions has been found: 
[counterConflicts=1, hashConflicts=0]
+Update counter conflicts:
+Conflict partition: PartitionKey [grpId=1544803905, grpName=default, partId=5]
+----
+
+The command compares partition update counters and partition hashes between 
primary and backup copies.
+Run it only when data updates are stopped.
+If updates are active, the command can report false conflicts because copies 
are changing while hashes are being calculated.

Review Comment:
   ```suggestion
   If updates are active, the command can report false positive conflicts 
because copies are changing while hashes are being calculated.
   ```



##########
docs/_docs/perf-and-troubleshooting/general-perf-tips.adoc:
##########
@@ -47,3 +47,458 @@ queries with JOINs at massive scale and expect significant 
performance benefits.
 
 * Adjust link:data-rebalancing[data rebalancing settings] to ensure that 
rebalancing completes faster when your cluster topology changes.
 
+== How to assess cluster health
+
+Cluster health is a complex thing. Apache Ignite is capable of demonstrating 
great performance across different scenarios with varying loads. Therefore, in 
general terms, a healthy cluster is one whose behavior aligns with your 
expectations. However, there are some universal aspects that apply to all 
deployments and warrant attention.

Review Comment:
   ```suggestion
   Cluster health is a complex set of metrics and states that need to be 
analyzed together. Apache Ignite is capable of demonstrating great performance 
across different scenarios with varying loads. Therefore, in general terms, a 
healthy cluster is one whose behavior aligns with your expectations. However, 
there are some universal aspects that apply to all deployments and warrant 
attention.
   ```



##########
docs/_docs/perf-and-troubleshooting/general-perf-tips.adoc:
##########
@@ -47,3 +47,458 @@ queries with JOINs at massive scale and expect significant 
performance benefits.
 
 * Adjust link:data-rebalancing[data rebalancing settings] to ensure that 
rebalancing completes faster when your cluster topology changes.
 
+== How to assess cluster health
+
+Cluster health is a complex thing. Apache Ignite is capable of demonstrating 
great performance across different scenarios with varying loads. Therefore, in 
general terms, a healthy cluster is one whose behavior aligns with your 
expectations. However, there are some universal aspects that apply to all 
deployments and warrant attention.
+
+It is important to understand that a healthy cluster may undergo planned 
topology changes or temporary load spikes.
+
+The key properties are:
+
+* The cluster is in the intended link:monitoring-metrics/cluster-states[state] 
and serves only the operations allowed by that state.
+* Baseline topology, when it is used or managed manually, matches the expected 
data-bearing server nodes.
+* Data remains consistent: 
link:tools/control-script#verifying-partition-checksums[`idle_verify`] reports 
no partition conflicts when the cluster is idle.
+* Expected nodes are present, and no unexpected repeated node `JOIN`, `LEFT`, 
or `FAIL` events or node segmentation are reported.
+* link:configuring-caches/partition-loss-policy[Lost partitions] are absent.
+* The number and age of 
link:tools/control-script#transaction-management[long-running transactions] do 
not keep increasing; active 
link:data-modeling/data-partitioning#partition-map-exchange[Partition Map 
Exchanges (PMEs)] complete; 
link:monitoring-metrics/new-metrics-system#monitoring-rebalancing[rebalancing] 
finishes; 
link:monitoring-metrics/new-metrics-system#monitoring-checkpointing-operations[checkpoint-related
 metrics] and link:monitoring-metrics/new-metrics#thread-pools[executor queue 
sizes] return to their usual ranges after the 
link:monitoring-metrics/new-metrics-system#monitoring-topology[topology] has 
stabilized and the application workload has returned to its expected level.
+There is no single command or metric that proves cluster health for every 
deployment.
+
+Use several signals together.
+A simple client connection or SQL liveness check can prove only that a 
particular client or query path is reachable; it does not check user 
partitions, backup consistency, baseline membership, or all server nodes.
+
+=== Check the Intended State and Node Membership
+
+Start with the link:monitoring-metrics/cluster-states[cluster state].
+
+Run:
+
+[source,shell]
+----
+control.(sh|bat) --state
+----
+
+Relevant output:
+
+[source,text]
+----
+Command [STATE] started
+Arguments: --state
+--------------------------------------------------------------------------------
+Cluster state: ACTIVE
+Command [STATE] finished with code: 0
+----
+
+`ACTIVE` is expected for normal read-write operation.
+`ACTIVE_READ_ONLY` is normal when read-only operation was intentionally 
enabled.
+`INACTIVE` is acceptable only when it matches the current operation, for 
example planned maintenance; an inactive cluster does not serve the data 
workload.
+The criterion is whether the actual state matches the state that was 
intentionally set for the deployment.
+
+Check baseline topology when the cluster uses persistence, when baseline 
autoadjustment is disabled, or when you intentionally manage the set of 
data-bearing nodes.
+In pure in-memory clusters with the default immediate autoadjustment, baseline 
topology normally follows the current server topology automatically.
+If autoadjustment is disabled, the baseline changes only after an operator 
changes it.
+If autoadjustment is configured with a non-zero timeout, the baseline is 
updated only after the topology remains unchanged for that timeout.
+In both cases, run `control.(sh|bat) --baseline` and compare `Baseline nodes` 
and `Other nodes` with the expected set of server nodes.
+
+Run:
+
+[source,shell]
+----
+control.(sh|bat) --baseline
+----
+
+Relevant output for a cluster where all baseline nodes are online:
+
+[source,text]
+----
+Cluster state: ACTIVE
+Current topology version: 3
+Baseline auto adjustment disabled: softTimeout=300000
+
+Current topology version: 3 (Coordinator: ConsistentId=node-1, Order=1)
+
+Baseline nodes:
+    ConsistentId=node-1, State=ONLINE, Order=1
+    ConsistentId=node-2, State=ONLINE, Order=2
+    ConsistentId=node-3, State=ONLINE, Order=3
+--------------------------------------------------------------------------------
+Number of baseline nodes: 3
+
+Other nodes not found.
+----
+
+Example: one baseline node is offline:
+
+[source,text]
+----
+Baseline nodes:
+    ConsistentId=node-1, State=ONLINE, Order=1
+    ConsistentId=node-2, State=OFFLINE, Order=2
+    ConsistentId=node-3, State=ONLINE, Order=3
+--------------------------------------------------------------------------------
+Number of baseline nodes: 3
+----
+
+If a baseline node is `OFFLINE`, an expected data-bearing server is missing.
+If other primary or backup copies are available, its absence does not cause 
partition loss.
+To check for partition loss, query the partition states as described in 
<<confirm-that-rebalancing-converges,Confirm That Rebalancing Converges>>.
+
+If an online server node has joined the cluster but is not in the baseline, 
the command shows it under `Other nodes`:
+
+[source,text]
+----
+Other nodes:
+    ConsistentId=node-4, Order=4
+Number of other nodes: 1
+----
+
+The baseline contains server nodes that are intended to store a data.
+Client nodes are not a part of the baseline.
+An online server node in `Other nodes` is not always an error: the node may 
have been prepared intentionally but not yet introduced into the data topology.
+If the node is expected to store data, first check the 
link:clustering/baseline-topology#baseline-topology-autoadjustment[baseline 
auto-adjustment policy] and the current maintenance or scale-out procedure, 
then use the documented baseline change procedure.
+Changing the baseline can start link:data-rebalancing[rebalancing]: partitions 
are redistributed according to the new affinity assignment.
+Plan for the additional network, CPU, and storage load, especially in clusters 
with persistence.
+
+Use topology changes to distinguish planned activity from instability.
+A server `JOIN`, `LEFT`, or `FAIL` event changes cluster membership and 
triggers link:data-modeling/data-partitioning#partition-map-exchange[Partition 
Map Exchanges (PME)].
+A dynamic cache start or stop can also trigger PME even when no server node 
joins, leaves, or fails.
+Therefore, a topology version or 
link:monitoring-metrics/new-metrics#partition-map-exchange[PME metric] PME 
metric change is useful only when interpreted together with maintenance actions 
and node logs.
+For example, these metrics should be used for historical distributions; there 
is no universal duration threshold.
+
+[source,sql]
+----
+SELECT NAME, VALUE
+FROM SYS.METRICS
+WHERE NAME IN (
+    'pme.Duration',
+    'pme.CacheOperationsBlockedDuration'
+)
+ORDER BY NAME;
+----
+
+Run `control.(sh|bat) --baseline` repeatedly or monitor topology metrics to 
confirm that membership is stable when no maintenance is in progress.
+In logs, look for repeated node join, left, fail, segmentation, and 
exchange-worker messages.
+Investigate unexpected repeated `JOIN`, `LEFT`, or `FAIL` events, node 
segmentation, network failures, or a PME that does not finish.
+For PME metrics and transaction checks, see 
<<check-transactions-and-sql-queries,Check Transactions and SQL Queries>>.
+
+=== Verify Partition Consistency
+
+When the cluster is expected to be idle, run:
+
+[source,shell]
+----
+control.(sh|bat) --cache idle_verify
+----
+
+Successful result:
+
+[source,text]
+----
+The check procedure has finished, no conflicts have been found.
+----
+
+The beginning of a conflict result uses this format:
+
+[source,text]
+----
+The check procedure has failed, conflict partitions has been found: 
[counterConflicts=1, hashConflicts=0]
+Update counter conflicts:
+Conflict partition: PartitionKey [grpId=1544803905, grpName=default, partId=5]
+----
+
+The command compares partition update counters and partition hashes between 
primary and backup copies.
+Run it only when data updates are stopped.
+If updates are active, the command can report false conflicts because copies 
are changing while hashes are being calculated.
+Partitions in `MOVING` or `LOST` state may be skipped, so the result can be 
incomplete.
+A successful `idle_verify` result is an important confirmation of consistency, 
but it still does not prove overall cluster health check success.
+
+[#confirm-that-rebalancing-converges]
+=== Confirm That Rebalancing Converges
+
+Immediately after an intended topology or cache event, `MOVING` and `RENTING` 
counts can be non-zero.
+After the topology becomes stable, repeat the query and confirm that both 
counts decrease and eventually disappear.
+A `LOST` count greater than zero is not a normal transient rebalance state.
+
+[source,sql]
+----
+SELECT STATE, COUNT(*) AS PARTITION_COUNT
+FROM SYS.PARTITION_STATES
+WHERE STATE IN ('MOVING', 'RENTING', 'LOST')
+GROUP BY STATE
+ORDER BY STATE;
+----
+
+It is also possible to track some node-local cache-group metrics, such as:
+
+* LocalNodeMovingPartitionsCount;
+* LocalNodeRentingPartitionsCount;
+* LocalNodeRentingEntriesCount.
+
+Use the 
link:monitoring-metrics/system-views#partition_states[PARTITION_STATES] system 
view to check partition states:
+
+* `OWNING`: the node is the current primary or backup owner.
+* `MOVING`: a partition copy is being loaded on the node during rebalance.
+* `RENTING`: an old copy is being removed after ownership changes.
+* `EVICTED`: the partition is absent on a node that is no longer an owner; 
this is not an error by itself.
+* `LOST`: the partition is unavailable and must not be used; investigate 
immediately.
+
+[source,sql]
+----
+SELECT CACHE_GROUP_ID, PARTITION_ID, NODE_ID, STATE, IS_PRIMARY
+FROM SYS.PARTITION_STATES
+WHERE STATE IN ('MOVING', 'RENTING', 'LOST')
+ORDER BY STATE, CACHE_GROUP_ID, PARTITION_ID, NODE_ID;
+----
+
+In steady state, this query usually should not return `MOVING`, `RENTING`, or 
`LOST` rows.
+`MOVING` and `RENTING` are expected right after an intended topology or cache 
event, but their count should decrease.
+`LOST` is not a normal transient state.
+
+Example: one partition is lost:
+
+[source,text]
+----
+CACHE_GROUP_ID | PARTITION_ID | NODE_ID                              | STATE | 
IS_PRIMARY
+1544803905     | 5            | 0f4d6f30-3e04-4f68-b6a2-6b89f1795c0d | LOST  | 
true
+----
+
+If the query returns `LOST`, follow the 
link:configuring-caches/partition-loss-policy[Partition Loss Policy] recovery 
procedure.
+If a failed node returns, its persistent data may become available again, but 
the affected partitions remain in the `LOST` state.
+If the required data is available again, reset the lost partitions.
+Before resetting lost partitions, ensure that at least one complete and 
up-to-date copy of each lost partition is available.
+
+For a persistent cluster, follow the 
link:configuring-caches/partition-loss-policy#clusters-with-persistence[recovery
 procedure for clusters with persistence].
+Return all nodes in the baseline topology before resetting lost partitions, or 
stop the cluster, start all nodes including the failed nodes, and activate the 
cluster.
+If some nodes cannot be returned, exclude them from the baseline topology only 
after determining whether their unavailable partition copies contain data that 
still has to be recovered.
+
+For example, assume that a partition has one primary and one backup copy on 
nodes A and B. If A leaves, B can continue accepting writes while it remains 
available.
+If B subsequently fails and only A returns, A can contain an older copy that 
does not include the writes accepted by B after A left.
+
+After restoring a complete and up-to-date partition copy, or after explicitly 
accepting that the unavailable updates cannot be recovered, reset the lost 
partitions:
+
+[source,shell]
+----
+control.(sh|bat) --cache reset_lost_partitions cacheName1,cacheName2,...
+----
+
+`reset_lost_partitions` only clears the `LOST` state. It does not reconstruct 
updates that are absent from every currently available copy.
+
+[#check-execution-queues]
+=== Check Execution Queues
+
+Ignite has several internal executors. A regular thread pool executes tasks 
from a shared queue. These queues may grow for a short time under load, but 
they should not grow continuously. Sustained queue growth means that a node is 
not keeping up with the workload or that message processing is impaired. The 
same logic applies to the striped executor.
+
+The striped executor divides internal cache and transaction tasks between 
independent stripes: tasks in the same stripe run sequentially, while different 
stripes can run in parallel.
+If a stripe is blocked, tasks related to that stripe can accumulate even when 
overall CPU usage does not look high.
+
+Check queue metrics on every server node:
+
+[source,sql]
+----
+SELECT NAME, VALUE
+FROM SYS.METRICS
+WHERE NAME IN ('io.communication.OutboundMessagesQueueSize'
+,'io.discovery.MessageWorkerQueueSize'
+,'threadPools.StripedExecutor.TotalQueueSize'
+,'threadPools.StripedExecutor.DetectStarvation'
+)
+   OR NAME LIKE 'threadPools.%.QueueSize'
+ORDER BY NAME;
+----
+
+Inspect queued striped tasks when the striped queue does not drain:
+
+[source,sql]
+----
+SELECT STRIPE_INDEX, THREAD_NAME, TASK_NAME, DESCRIPTION
+FROM SYS.STRIPED_THREADPOOL_QUEUE
+ORDER BY STRIPE_INDEX, THREAD_NAME;
+----
+
+As mentioned above, short non-zero queues are acceptable under load.
+
+On an idle node, queues usually return to zero.
+Investigate continuous growth, lack of drain after load stops, repeated 
`DetectStarvation=true`, or repeated starvation warnings in logs.
+There is no universal absolute threshold.
+Queue metrics are node-local, so collect them from all server nodes.
+
+JMX uses the metric registry name to build `group` and `name` in the MBean 
object name.
+The following mappings are useful for queue checks:
+
+* Registry `io.communication` is exposed as JMX group `io`, bean name 
`communication`; the attribute is `OutboundMessagesQueueSize`.
+* Registry `io.discovery` is exposed as JMX group `io`, bean name `discovery`; 
the attribute is `MessageWorkerQueueSize`.
+* Registry `threadPools.StripedExecutor` is exposed as JMX group 
`threadPools`, bean name `StripedExecutor`; the attributes include 
`TotalQueueSize`, `StripesQueueSizes`, and `DetectStarvation`.
+* Regular pools such as `threadPools.GridSystemExecutor` expose `QueueSize`.
+
+For JMX object names and SQL metric access, see 
link:monitoring-metrics/new-metrics-system#jmx[JMX] and 
link:monitoring-metrics/new-metrics-system#sql-view[SQL View].
+
+.JConsole view of node-local striped executor queue metrics
+image::perf-and-troubleshooting/images/healthy-cluster-queues-jconsole.png[JConsole
 MBeans view showing threadPools/StripedExecutor and queue-related attributes]
+
+[#check-transactions-and-sql-queries]
+=== Check Transactions and SQL Queries
+
+A transaction or query is not unhealthy merely because it runs for some time.
+Investigate when the number or age of active operations continues to increase 
after the load drops, or when the same operations repeatedly block other work.
+
+Use the transaction command to list long transactions:
+
+[source,shell]
+----
+control.(sh|bat) --tx --min-duration 60 --servers --order DURATION
+----
+
+The value `60` is only an example diagnostic filter in seconds, not a 
universal production threshold.
+
+Use the system views for current transactions and SQL queries:
+
+[source,sql]
+----
+SELECT XID, STATE, START_TIME, DURATION, KEYS_COUNT, LABEL
+FROM SYS.TRANSACTIONS
+ORDER BY DURATION DESC;
+----
+
+[source,sql]
+----
+SELECT QUERY_ID, START_TIME, DURATION, INITIATOR_ID, SQL
+FROM SYS.SQL_QUERIES
+ORDER BY DURATION DESC;
+----
+
+Track related metrics:
+
+[source,sql]
+----
+SELECT NAME, VALUE
+FROM SYS.METRICS
+WHERE NAME IN (
+    'tx.OwnerTransactionsNumber',
+    'tx.TransactionsHoldingLockNumber',
+    'tx.LockedKeysNumber',
+    'pme.Duration',
+    'pme.CacheOperationsBlockedDuration'
+)
+ORDER BY NAME;
+----
+
+Non-zero transaction counters are normal while work is running.
+The problem is sustained growth, increasing age of the oldest operations, and 
failure to return to the usual range after the workload drops.
+For view definitions and metrics, see 
link:monitoring-metrics/system-views#transactions[TRANSACTIONS], 
link:monitoring-metrics/system-views#sql_queries[SQL_QUERIES], 
link:monitoring-metrics/new-metrics#transactions[transaction metrics], and 
link:data-modeling/data-partitioning#partition-map-exchange[Partition Map 
Exchange metrics].
+
+[#check-partition-map-exchange]
+llink:data-modeling/data-partitioning#partition-map-exchange[Partition Map 
Exchanges (PMEs)] synchronizes partition distribution after topology and cache 
changes.
+At one stage, PME waits for incomplete transactions to finish.
+A long transaction can delay a node join, cache start, and other operations 
that depend on exchange.
+
+`TransactionConfiguration.setTxTimeoutOnPartitionMapExchange(...)` is 
described in 
link:key-value-api/transactions#long-running-transactions-termination[Long 
Running Transactions Termination].
+The default is `0`, which means transactions are not rolled back because of a 
PME timeout.
+The timeout is applied only when PME starts.
+Incomplete transactions that exceed the configured value can be rolled back.
+Applications must handle `TransactionRollbackException` and retry where 
appropriate; see 
link:key-value-api/transactions#handling-failed-transactions[Handling Failed 
Transactions].
+Do not use a universal timeout value.
+Choose a value above the normal duration of legitimate transactions with a 
justified safety margin, and test application behavior when rollback happens.
+
+[#check-checkpoint-pressure]
+=== Check Checkpoint Pressure When Persistence Is Enabled
+
+This check applies only to data regions with Native Persistence enabled.
+A pure in-memory cluster does not perform persistence checkpoints for its 
in-memory regions.
+
+A link:persistence/native-persistence#checkpointing[checkpoint] writes dirty 
pages from RAM to partition files.
+Checkpointing itself is a normal background operation.
+The problem starts when the application write rate exceeds the effective 
storage write speed.
+Under checkpoint-buffer or dirty-page pressure, Ignite can throttle update 
threads.
+If the checkpoint buffer is exhausted, update processing can stop until the 
checkpoint completes.
+
+Monitor these metrics for persistent data regions and data storage:
+
+* `io.dataregion.<region>.DirtyPages`
+* `io.dataregion.<region>.CheckpointBufferSize`
+* `io.dataregion.<region>.UsedCheckpointBufferSize`
+* `io.dataregion.<region>.TotalThrottlingTime`
+* `io.datastorage.LastCheckpointStart`
+* `io.datastorage.LastCheckpointDuration`
+* `io.datastorage.LastCheckpointPagesWriteDuration`
+* `io.datastorage.LastCheckpointTotalPagesNumber`
+* `io.datastorage.LastCheckpointFsyncDuration`
+
+It looks like this in ignite.log

Review Comment:
   ```suggestion
   Corresponding records in <SOME_formatting ? bold ? italic ?>ignite.log looks 
like as follows:
   ```



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


Reply via email to