On 02-04-2026 01:46 pm, Raag Jadav wrote:
On Wed, Mar 18, 2026 at 12:10:20PM +0530, Mallesh Koujalagi wrote:
Add documentation for the DRM_WEDGE_RECOVERY_COLD_RESET recovery
method introduced for handling power management unit errors. This method is
designated for severe errors that compromise core device functionality
and are unrecoverable via recovery mechanisms such as driver reload or PCIe
bus reset. The documentation clarifies when this recovery method should be
used and its implications for userspace applications.
Aesthetic nit: We usually try to utilize the full 75 character space where
possible (in all patches).
v2:
- Add several instead of number to avoid update. (Jani)
Signed-off-by: Mallesh Koujalagi <[email protected]>
---
Documentation/gpu/drm-uapi.rst | 73 +++++++++++++++++++++++++++++++++-
1 file changed, 72 insertions(+), 1 deletion(-)
diff --git a/Documentation/gpu/drm-uapi.rst b/Documentation/gpu/drm-uapi.rst
index d98428a592f1..5b63f1c17b9b 100644
--- a/Documentation/gpu/drm-uapi.rst
+++ b/Documentation/gpu/drm-uapi.rst
@@ -418,7 +418,7 @@ needed.
Recovery
--------
-Current implementation defines four recovery methods, out of which, drivers
+Current implementation defines several recovery methods, out of which, drivers
can use any one, multiple or none. Method(s) of choice will be sent in the
uevent environment as ``WEDGED=<method1>[,..,<methodN>]`` in order of less to
more side-effects. See the section `Vendor Specific Recovery`_
@@ -435,6 +435,7 @@ following expectations.
rebind unbind + bind driver
bus-reset unbind + bus reset/re-enumeration + bind
vendor-specific vendor specific recovery method
+ cold-reset full device cold reset required
Does this work without unbind + bind?
No. In cold-reset method, we do unbind + power cycle + rescan/bind.
Thanks,
-/Mallesh
unknown consumer policy
=============== ========================================
@@ -446,6 +447,27 @@ telemetry information (devcoredump, syslog). This is useful because the first
hang is usually the most critical one which can result in consequential hangs
or
complete wedging.
+Cold Reset Recovery
+-------------------
+
+The ``WEDGED=cold-reset`` event indicates that the device has encountered
+power management unit errors that affect core functionality that cannot be
+resolved through recovery mechanisms.
Here cold-reset itself is introduced as a recovery method, so this could
use better phrasing. Also, we try to be consistent with terminologies
already in place so that the document is easy to follow.
+This recovery method is reserved for power management unit error conditions
where the
+device state cannot be restored via:
+
+- Driver unbind/rebind operations
+- PCIe bus reset and re-enumeration
+- Device Function Level Reset (FLR)
+- Warm device resets
These are already documented so not sure if it's worth repeating them.
+Such power management unit error state typically persists across all
software-based
+recovery attempts. Only a complete device power cycle can restore
+normal operation.
+
+Upon receiving a ``WEDGED=cold-reset`` event, userspace should initiate
+a full cold reset of the affected device to restore functionality.
This is already covered in consumer expectations so can be dropped.
AI assistance is useful at times but we'd want to make sure that things
remain objective and translatable ;)
Raag
Vendor Specific Recovery
------------------------
@@ -524,6 +546,55 @@ Recovery script::
echo -n $DEVICE > $DRIVER/unbind
echo -n $DEVICE > $DRIVER/bind
+Example - cold-reset
+--------------------
+
+Udev rule::
+
+ SUBSYSTEM=="drm", ENV{WEDGED}=="cold-reset", DEVPATH=="*/drm/card[0-9]",
+ RUN+="/path/to/cold-reset.sh $env{DEVPATH}"
+
+Recovery script::
+
+ #!/bin/sh
+
+ [ -z "$1" ] && echo "Usage: $0 <device-path>" && exit 1
+
+ # Get device
+ DEVPATH=$(readlink -f /sys/$1/device 2>/dev/null || readlink -f /sys/$1)
+ DEVICE=$(basename $DEVPATH)
+
+ echo "Cold reset: $DEVICE"
+
+ # Try slot power reset first
+ SLOT=$(find /sys/bus/pci/slots/ -type l 2>/dev/null | while read slot; do
+ ADDR=$(cat "$slot" 2>/dev/null)
+ [ -n "$ADDR" ] && echo "$DEVICE" | grep -q "^$ADDR" && basename $(dirname
"$slot") && break
+ done)
+
+ if [ -n "$SLOT" ]; then
+ echo "Using slot $SLOT"
+
+ # Unbind driver
+ [ -e "/sys/bus/pci/devices/$DEVICE/driver" ] && \
+ echo "$DEVICE" > /sys/bus/pci/devices/$DEVICE/driver/unbind 2>/dev/null
+
+ # Remove device
+ echo 1 > /sys/bus/pci/devices/$DEVICE/remove
+
+ # Power cycle slot
+ echo 0 > /sys/bus/pci/slots/$SLOT/power
+ sleep 2
+ echo 1 > /sys/bus/pci/slots/$SLOT/power
+ sleep 1
+
+ # Rescan
+ echo 1 > /sys/bus/pci/rescan
+ echo "Done!"
+ else
+ echo "No slot found"
+ fi
+
Customization
-------------
--
2.34.1