When PUNIT (power management unit) errors are detected that persist across warm resets, mark the device as wedged with DRM_WEDGE_RECOVERY_COLD_RESET and notify userspace that a complete device power cycle is required to restore normal operation.
Signed-off-by: Mallesh Koujalagi <[email protected]> Reviewed-by: Raag Jadav <[email protected]> --- v3: - Use PUNIT instead of PMU. (Riana) - Use consistent wording. - Remove log. (Raag) v4: - Make function static. (Raag) v5: - Remove kdoc for static function. (Raag) - Remove xe_ prefix for static function. v9: - Remove unwanted header. (Sashiko) --- drivers/gpu/drm/xe/xe_ras.c | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c index a31e06b8aa67..92b4181026cb 100644 --- a/drivers/gpu/drm/xe/xe_ras.c +++ b/drivers/gpu/drm/xe/xe_ras.c @@ -236,6 +236,12 @@ static u8 handle_core_compute_errors(struct xe_ras_error_array *arr) return XE_RAS_RECOVERY_ACTION_RECOVERED; } +static void punit_error_handler(struct xe_device *xe) +{ + xe_device_set_wedged_method(xe, DRM_WEDGE_RECOVERY_COLD_RESET); + xe_device_declare_wedged(xe); +} + static u8 handle_soc_internal_errors(struct xe_device *xe, struct xe_ras_error_array *arr) { struct xe_ras_soc_error *info = (void *)arr->details; @@ -267,7 +273,7 @@ static u8 handle_soc_internal_errors(struct xe_device *xe, struct xe_ras_error_a xe_err(xe, "[RAS]: PUNIT %s detected: 0x%x\n", sev_to_str(counter->common.severity), ieh_error->global_error_status); - /* TODO: Add PUNIT error handling */ + punit_error_handler(xe); return XE_RAS_RECOVERY_ACTION_DISCONNECT; } } -- 2.48.1
