alex-plekhanov commented on code in PR #13339:
URL: https://github.com/apache/ignite/pull/13339#discussion_r4030366511
##########
docs/_docs/key-value-api/transactions.adoc:
##########
@@ -358,43 +357,23 @@ TcpDiscoveryNode [id=<id>, addrs=[<address>], order=1,
ver=2.18.0#YYYYMMDD-sha1:
Tx: [xid=<id>, label=null, state=ACTIVE, startTime=YYYY-MM-DD
15:15:16.852, duration=3 sec, isolation=REPEATABLE_READ,
concurrency=PESSIMISTIC, topVer=AffinityTopologyVersion [topVer=3,
minorTopVer=2], timeout=0 sec, size=100, dhtNodes=[<id>, <id>], nearXid=<id>,
parentNodeIds=[<id>]]
Command [TX] finished with code: 0
----
-. Extract the `xid` from all transactions across all nodes and remove
duplicate `xid` values. Duplicates occur because the same transaction can exist
on multiple nodes — in this case, its `xid` will appear in the command output
multiple times.
-. For each `xid`, run the `tx kill` command to roll back that transaction.
Example call: `control.sh --tx --xid <id> --kill`.
-
-* Second termination stage (optional) — restarting nodes that block
transaction completion. Check if the first stage was sufficient to handle LRT:
-. After completing the first stage, wait for a period equal to the configured
transaction timeout.
-. Run the `control.sh --tx --min-duration <transaction_timeout>` command.
-If the list is empty, all LRTs have been successfully canceled, and the LRT
impact should be resolved.
-If the list is not empty, the command output will contain information about
nodes and the transactions running on them:
-+
+To cancel the transactions, whose execution time takes longer than expected,
use the `control.sh --tx --min-duration <transaction_duration_seconds> --kill`
command. For example, to cancel the transactions that have been running for
more than 100 seconds, execute the following command:
+[source, shell]
----
-Command [TX] started
-Arguments: --tx
---------------------------------------------------------------------------------
-Matching transactions:
-TcpDiscoveryNode [id=<id>, addrs=[<address>], order=1,
ver=2.18.0#YYYYMMDD-sha1:00000000, isClient=false,
consistentId=gridCommandHandlerTest0]
- Tx: [xid=<id>, label=null, state=ACTIVE, startTime=YYYY-MM-DD
16:58:45.348, duration=3 sec, isolation=REPEATABLE_READ,
concurrency=PESSIMISTIC, topVer=AffinityTopologyVersion [topVer=3,
minorTopVer=2], timeout=0 sec, size=100, dhtNodes=[<id>, <id>], nearXid=<id>,
parentNodeIds=[<id>]]
- Tx: [xid=<id>, label=label1, state=ACTIVE, startTime=YYYY-MM-DD
16:58:48.451, duration=0 sec, isolation=READ_COMMITTED,
concurrency=PESSIMISTIC, topVer=AffinityTopologyVersion [topVer=3,
minorTopVer=2], timeout=2147483 sec, size=111, dhtNodes=[<id>, <id>],
nearXid=<id>, parentNodeIds=[<id>]]
-TcpDiscoveryNode [id=<id>, addrs=[<address>], order=2,
ver=2.18.0#YYYYMMDD-sha1:00000000, isClient=false,
consistentId=gridCommandHandlerTest1]
- Tx: [xid=<id>, label=null, state=ACTIVE, startTime=YYYY-MM-DD
16:58:45.348, duration=3 sec, isolation=REPEATABLE_READ,
concurrency=PESSIMISTIC, topVer=AffinityTopologyVersion [topVer=3,
minorTopVer=2], timeout=0 sec, size=1, dhtNodes=[<id>], nearXid=<id>,
parentNodeIds=[<id>]]
-TcpDiscoveryNode [id=<id>, addrs=[<address>], order=3,
ver=2.18.0#YYYYMMDD-sha1:00000000, isClient=true, consistentId=client]
- Tx: [xid=<id>, label=label2, state=PREPARING, startTime=YYYY-MM-DD
16:58:45.348, duration=3 sec, isolation=READ_COMMITTED, concurrency=OPTIMISTIC,
topVer=AffinityTopologyVersion [topVer=3, minorTopVer=2], timeout=0 sec,
size=11, dhtNodes=[<id>, <id>], nearXid=<id>, parentNodeIds=[<id>]]
-Command [TX] finished with code: 0
+control.sh --tx --min-duration 100 --kill
----
-. For each transaction, extract the value of the `dhtNodes` field.
-. For each ID in `dhtNodes`, obtain detailed node information using the
`control.sh --system-view NODES` command.
-. In the command output, find the node whose `nodeId` starts with the string
from `dhtNodes`.
-. Restart this node to unblock the transaction.
== Causes of LRT
* High system load, which leads to slower transaction processing. Indicators:
** log messages about the start of page eviction for in-memory data regions or
page replacement for persistent data regions;
** growth of queues in the striped pool;
** increased checkpoint duration for clusters with persistence enabled;
+** high cpu or disk utilization;
+** increased latency of cache operation;
** long GC pauses.
* Unstable network operation. Indicators: log messages about connection loss
between nodes, socket closures, or network timeouts triggering.
-* Resource-intensive operations (for example, deleting a cache that is part of
a cache group). Indicators: queue growth on individual threads of the system
pool and striped pool, while other threads and overall node utilization may
remain low.
-* Internal Apache Ignite bugs that can lead to deadlocks. Indicators: thread
dumps (`thread dump`) show threads waiting to acquire a read lock (`readLock`)
or write lock (`writeLock`) during checkpoint creation; over time, the thread
state does not change (stack trace remains unchanged).
+* Resource-intensive operations (for example, deleting a cache that is part of
a cache group, snapshots execution, index rebuild, rebalancing, etc.).
+* Deadlocks during transactions execution.
Review Comment:
`Deadlocks during transactions execution` can be treat as keys deadlocks,
but here we mean java-level deadlocks. So, maybe it's worth to leave `Internal
Apache Ignite bugs or java-level deadlocks in Apache Ignite code.' There is
some deadlock detection mechanism in thread dump, and it can detect some kind
of deadlocks. But unfortunately not all types (for example, deadlocks caused by
pool starvation can't be detected by thread dump). But we have also pool
starvation detector for some pools. So, it also can be mentioned here.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]