AMD General

Cannot send command during GPU reset, which may cause a system exception.

Best Regards,
Thomas
-----Original Message-----
From: Sun, Ce(Overlord) <[email protected]>
Sent: Friday, July 24, 2026 7:59 AM
To: [email protected]
Cc: Zhang, Hawking <[email protected]>; Chai, Thomas <[email protected]>; 
Zhou1, Tao <[email protected]>; Yang, Stanley <[email protected]>; Lazar, 
Lijo <[email protected]>; Sun, Ce(Overlord) <[email protected]>
Subject: [PATCH] drm/amd/ras: eliminate gpu reset lock protection for ras ta 
runtime command path

When PMFW enters mode-1 state, it will not respond to incoming messages.
Holding gpu reset lock before dispatching messages to PMFW would defer GPU 
recovery processing, cause undesirable timing skew.

Remove the gpu reset lock acquire and release logic inside 
send_ras_ta_runtime_cmd, prioritize GPU recovery execution over PMFW message 
delivery to fix timing dependency problems.

Signed-off-by: Ce Sun <[email protected]>
---
 drivers/gpu/drm/amd/ras/rascore/ras_psp.c | 6 ------
 1 file changed, 6 deletions(-)

diff --git a/drivers/gpu/drm/amd/ras/rascore/ras_psp.c 
b/drivers/gpu/drm/amd/ras/rascore/ras_psp.c
index 358f602b167d..fbe91eb67c0e 100644
--- a/drivers/gpu/drm/amd/ras/rascore/ras_psp.c
+++ b/drivers/gpu/drm/amd/ras/rascore/ras_psp.c
@@ -374,11 +374,6 @@ static int send_ras_ta_runtime_cmd(struct ras_core_context 
*ras_core,
        if (!cmd_mem)
                return -ENOMEM;

-       if (!ras_core_down_trylock_gpu_reset_lock(ras_core)) {
-               ret = -EACCES;
-               goto out;
-       }
-
        ras_cmd = (struct ras_ta_cmd *)cmd_mem->mem_cpu_addr;

        mutex_lock(&ta_ctx->ta_mutex);
@@ -422,7 +417,6 @@ static int send_ras_ta_runtime_cmd(struct ras_core_context 
*ras_core,

 unlock:
        mutex_unlock(&ta_ctx->ta_mutex);
-       ras_core_up_gpu_reset_lock(ras_core);
 out:
        ras_psp_put_gpu_mem(ras_core, cmd_mem);
        return ret;
--
2.34.1

Reply via email to