AMD General


Get Outlook for Android<https://aka.ms/AAb9ysg>

________________________________
From: Lazar, Lijo <[email protected]>
Sent: Tuesday, July 14, 2026 3:14:34 PM
To: SHANMUGAM, SRINIVASAN <[email protected]>; Deucher, Alexander 
<[email protected]>; [email protected] 
<[email protected]>
Cc: Liang, Prike <[email protected]>; Khatri, Sunil <[email protected]>
Subject: Re: [PATCH] drm/amdgpu/userq: properly account for resets



On 14-Jul-26 10:16 AM, SHANMUGAM, SRINIVASAN wrote:
> AMD General
>
>> -----Original Message-----
>> From: Alex Deucher <[email protected]>
>> Sent: Tuesday, July 14, 2026 2:09 AM
>> To: [email protected]
>> Cc: Deucher, Alexander <[email protected]>; SHANMUGAM,
>> SRINIVASAN <[email protected]>; Liang, Prike
>> <[email protected]>; Khatri, Sunil <[email protected]>
>> Subject: [PATCH] drm/amdgpu/userq: properly account for resets
>>
>> We need to increment the reset counter, force fence completion, and set the
>> wedged event when a user queue is reset.
>>
>> mes_userq_reset_queue() handles this for collateral damage, but the caller 
>> needs
>> to handle this directly for the original guilty queue.
>>
>> Signed-off-by: Alex Deucher <[email protected]>
>> Cc: Srinivasan Shanmugam <[email protected]>
>> Cc: Prike Liang <[email protected]>
>> Cc: Sunil Khatri <[email protected]>
>> ---
>>   drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 7 ++++++-
>>   1 file changed, 6 insertions(+), 1 deletion(-)
>>
>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>> b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>> index 6aa75da27f912..5e1262636e1e9 100644
>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>> @@ -146,8 +146,13 @@ static void amdgpu_userq_hang_detect_work(struct
>> work_struct *work)
>>                                                         queue, NULL, NULL);
>>                else
>>                        r = userq_funcs->reset(queue);
>> -             if (r)
>> +             if (r) {
>>                        gpu_reset = true;
>> +             } else {
>> +                     atomic_inc(&adev->gpu_reset_counter);
>> +                     amdgpu_userq_fence_driver_force_completion(queue);
>> +                     drm_dev_wedged_event(adev_to_drm(adev),
>> DRM_WEDGE_RECOVERY_NONE, NULL);
>> +             }
>>        } else {
>>                gpu_reset = true;
>>        }
>
> After the original queue was reset successfully, it did not update 
> gpu_reset_counter, complete its pending fences, or send the wedged event.
> mes_userq_reset_queue() already updates gpu_reset_counter, completes the 
> pending fences, and sends the wedged event for the other affected queues,
> but skips the original queue because it has already been reset.

What is the rationale of sending multiple device wedged events on a per
queue basis?

The question of whether drm_dev_wedged_event() should be emitted once per queue 
or once per overall recovery seems like a broader design discussion.

Thanks,
Srini



Reply via email to