Hi Riana,
On 30-03-2026 10:30 am, Tauro, Riana wrote:
On 3/18/2026 12:10 PM, Mallesh Koujalagi wrote:
Add documentation for the DRM_WEDGE_RECOVERY_COLD_RESET recovery
method introduced for handling power management unit errors. This
method is
designated for severe errors that compromise core device functionality
and are unrecoverable via recovery mechanisms such as driver reload
or PCIe
bus reset. The documentation clarifies when this recovery method
should be
used and its implications for userspace applications.
v2:
- Add several instead of number to avoid update. (Jani)
Signed-off-by: Mallesh Koujalagi <[email protected]>
---
Documentation/gpu/drm-uapi.rst | 73 +++++++++++++++++++++++++++++++++-
1 file changed, 72 insertions(+), 1 deletion(-)
diff --git a/Documentation/gpu/drm-uapi.rst
b/Documentation/gpu/drm-uapi.rst
index d98428a592f1..5b63f1c17b9b 100644
--- a/Documentation/gpu/drm-uapi.rst
+++ b/Documentation/gpu/drm-uapi.rst
@@ -418,7 +418,7 @@ needed.
Recovery
--------
-Current implementation defines four recovery methods, out of
which, drivers
+Current implementation defines several recovery methods, out of
which, drivers
can use any one, multiple or none. Method(s) of choice will be sent
in the
uevent environment as ``WEDGED=<method1>[,..,<methodN>]`` in order
of less to
more side-effects. See the section `Vendor Specific Recovery`_
@@ -435,6 +435,7 @@ following expectations.
rebind unbind + bind driver
bus-reset unbind + bus reset/re-enumeration + bind
vendor-specific vendor specific recovery method
+ cold-reset full device cold reset required
unknown consumer policy
=============== ========================================
@@ -446,6 +447,27 @@ telemetry information (devcoredump, syslog).
This is useful because the first
hang is usually the most critical one which can result in
consequential hangs or
complete wedging.
+Cold Reset Recovery
+-------------------
+
+The ``WEDGED=cold-reset`` event indicates that the device has
encountered
+power management unit errors that affect core functionality that
cannot be
Power management errors may be only xe usecase. Keep the
documentation vendor-agnostic.
Some vendors may want to use it for a different usecase
+resolved through recovery mechanisms.
+
+This recovery method is reserved for power management unit error
conditions where the
Same as above
Agree, will rework!
Thanks,
-/Mallesh
Thanks
Riana
+device state cannot be restored via:
+
+- Driver unbind/rebind operations
+- PCIe bus reset and re-enumeration
+- Device Function Level Reset (FLR)
+- Warm device resets
+
+Such power management unit error state typically persists across all
software-based
+recovery attempts. Only a complete device power cycle can restore
+normal operation.
+
+Upon receiving a ``WEDGED=cold-reset`` event, userspace should initiate
+a full cold reset of the affected device to restore functionality.
Vendor Specific Recovery
------------------------
@@ -524,6 +546,55 @@ Recovery script::
echo -n $DEVICE > $DRIVER/unbind
echo -n $DEVICE > $DRIVER/bind
+Example - cold-reset
+--------------------
+
+Udev rule::
+
+ SUBSYSTEM=="drm", ENV{WEDGED}=="cold-reset",
DEVPATH=="*/drm/card[0-9]",
+ RUN+="/path/to/cold-reset.sh $env{DEVPATH}"
+
+Recovery script::
+
+ #!/bin/sh
+
+ [ -z "$1" ] && echo "Usage: $0 <device-path>" && exit 1
+
+ # Get device
+ DEVPATH=$(readlink -f /sys/$1/device 2>/dev/null || readlink -f
/sys/$1)
+ DEVICE=$(basename $DEVPATH)
+
+ echo "Cold reset: $DEVICE"
+
+ # Try slot power reset first
+ SLOT=$(find /sys/bus/pci/slots/ -type l 2>/dev/null | while read
slot; do
+ ADDR=$(cat "$slot" 2>/dev/null)
+ [ -n "$ADDR" ] && echo "$DEVICE" | grep -q "^$ADDR" &&
basename $(dirname "$slot") && break
+ done)
+
+ if [ -n "$SLOT" ]; then
+ echo "Using slot $SLOT"
+
+ # Unbind driver
+ [ -e "/sys/bus/pci/devices/$DEVICE/driver" ] && \
+ echo "$DEVICE" > /sys/bus/pci/devices/$DEVICE/driver/unbind
2>/dev/null
+
+ # Remove device
+ echo 1 > /sys/bus/pci/devices/$DEVICE/remove
+
+ # Power cycle slot
+ echo 0 > /sys/bus/pci/slots/$SLOT/power
+ sleep 2
+ echo 1 > /sys/bus/pci/slots/$SLOT/power
+ sleep 1
+
+ # Rescan
+ echo 1 > /sys/bus/pci/rescan
+ echo "Done!"
+ else
+ echo "No slot found"
+ fi
+
Customization
-------------