On 14-Jul-26 3:57 PM, SHANMUGAM, SRINIVASAN wrote:
AMD General




Get Outlook for Android <https://aka.ms/AAb9ysg>

------------------------------------------------------------------------
*From:* Lazar, Lijo <[email protected]>
*Sent:* Tuesday, July 14, 2026 3:14:34 PM
*To:* SHANMUGAM, SRINIVASAN <[email protected]>; Deucher, Alexander <[email protected]>; [email protected] <[email protected]> *Cc:* Liang, Prike <[email protected]>; Khatri, Sunil <[email protected]>
*Subject:* Re: [PATCH] drm/amdgpu/userq: properly account for resets



On 14-Jul-26 10:16 AM, SHANMUGAM, SRINIVASAN wrote:
 > AMD General
 >
 >> -----Original Message-----
 >> From: Alex Deucher <[email protected]>
 >> Sent: Tuesday, July 14, 2026 2:09 AM
 >> To: [email protected]
 >> Cc: Deucher, Alexander <[email protected]>; SHANMUGAM,
 >> SRINIVASAN <[email protected]>; Liang, Prike
 >> <[email protected]>; Khatri, Sunil <[email protected]>
 >> Subject: [PATCH] drm/amdgpu/userq: properly account for resets
 >>
>> We need to increment the reset counter, force fence completion, and set the
 >> wedged event when a user queue is reset.
 >>
>> mes_userq_reset_queue() handles this for collateral damage, but the caller needs
 >> to handle this directly for the original guilty queue.
 >>
 >> Signed-off-by: Alex Deucher <[email protected]>
 >> Cc: Srinivasan Shanmugam <[email protected]>
 >> Cc: Prike Liang <[email protected]>
 >> Cc: Sunil Khatri <[email protected]>
 >> ---
 >>   drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 7 ++++++-
 >>   1 file changed, 6 insertions(+), 1 deletion(-)
 >>
 >> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
 >> b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
 >> index 6aa75da27f912..5e1262636e1e9 100644
 >> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
 >> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
 >> @@ -146,8 +146,13 @@ static void amdgpu_userq_hang_detect_work(struct
 >> work_struct *work)
>>                                                         queue, NULL, NULL);
 >>                else
 >>                        r = userq_funcs->reset(queue);
 >> -             if (r)
 >> +             if (r) {
 >>                        gpu_reset = true;
 >> +             } else {
 >> +                     atomic_inc(&adev->gpu_reset_counter);
 >> +                     amdgpu_userq_fence_driver_force_completion(queue);
 >> +                     drm_dev_wedged_event(adev_to_drm(adev),
 >> DRM_WEDGE_RECOVERY_NONE, NULL);
 >> +             }
 >>        } else {
 >>                gpu_reset = true;
 >>        }
 >
> After the original queue was reset successfully, it did not update gpu_reset_counter, complete its pending fences, or send the wedged event. > mes_userq_reset_queue() already updates gpu_reset_counter, completes the pending fences, and sends the wedged event for the other affected queues,
 > but skips the original queue because it has already been reset.

What is the rationale of sending multiple device wedged events on a per
queue basis?

The question of whether drm_dev_wedged_event() should be emitted once per queue or once per overall recovery seems like a broader design discussion.


Along with that, also need to consider if device reset_counter needs to be incremented on a per queue basis or based on reset event recovery. It could get incremented multiple times inside this - mes_userq_reset_queue.

Thanks,
Lijo

Thanks,
Srini




Reply via email to