On Tue, Jul 14, 2026 at 10:10 AM SHANMUGAM, SRINIVASAN
<[email protected]> wrote:
>
> AMD General
>
> > -----Original Message-----
> > From: Alex Deucher <[email protected]>
> > Sent: Tuesday, July 14, 2026 7:26 PM
> > To: SHANMUGAM, SRINIVASAN <[email protected]>
> > Cc: Lazar, Lijo <[email protected]>; Deucher, Alexander
> > <[email protected]>; [email protected]; Liang, Prike
> > <[email protected]>; Khatri, Sunil <[email protected]>
> > Subject: Re: [PATCH] drm/amdgpu/userq: properly account for resets
> >
> > On Tue, Jul 14, 2026 at 9:49 AM SHANMUGAM, SRINIVASAN
> > <[email protected]> wrote:
> > >
> > > AMD General
> > >
> > > > -----Original Message-----
> > > > From: Lazar, Lijo <[email protected]>
> > > > Sent: Tuesday, July 14, 2026 4:02 PM
> > > > To: SHANMUGAM, SRINIVASAN <[email protected]>;
> > Deucher,
> > > > Alexander <[email protected]>; amd-
> > > > [email protected]
> > > > Cc: Liang, Prike <[email protected]>; Khatri, Sunil
> > > > <[email protected]>
> > > > Subject: Re: [PATCH] drm/amdgpu/userq: properly account for resets
> > > >
> > > >
> > > >
> > > > On 14-Jul-26 3:57 PM, SHANMUGAM, SRINIVASAN wrote:
> > > > > AMD General
> > > > >
> > > > >
> > > > >
> > > > >
> > > > > Get Outlook for Android <https://aka.ms/AAb9ysg>
> > > > >
> > > > > ------------------------------------------------------------------
> > > > > ----
> > > > > --
> > > > > *From:* Lazar, Lijo <[email protected]>
> > > > > *Sent:* Tuesday, July 14, 2026 3:14:34 PM
> > > > > *To:* SHANMUGAM, SRINIVASAN
> > <[email protected]>;
> > > > Deucher,
> > > > > Alexander <[email protected]>;
> > > > > [email protected] <[email protected]>
> > > > > *Cc:* Liang, Prike <[email protected]>; Khatri, Sunil
> > > > > <[email protected]>
> > > > > *Subject:* Re: [PATCH] drm/amdgpu/userq: properly account for
> > > > > resets
> > > > >
> > > > >
> > > > >
> > > > > On 14-Jul-26 10:16 AM, SHANMUGAM, SRINIVASAN wrote:
> > > > >  > AMD General
> > > > >  >
> > > > >  >> -----Original Message-----
> > > > >  >> From: Alex Deucher <[email protected]>  >> Sent:
> > > > > Tuesday, July 14, 2026 2:09 AM  >> To: [email protected]  
> > > > > >>
> > Cc:
> > > > > Deucher, Alexander <[email protected]>; SHANMUGAM,  >>
> > > > > SRINIVASAN <[email protected]>; Liang, Prike  >>
> > > > > <[email protected]>; Khatri, Sunil <[email protected]>  >>
> > > > > Subject: [PATCH] drm/amdgpu/userq: properly account for resets  >>
> > > > > >> We need to increment the reset counter, force fence completion,
> > > > > and set the  >> wedged event when a user queue is reset.
> > > > >  >>
> > > > >  >> mes_userq_reset_queue() handles this for collateral damage,
> > > > > but the caller needs  >> to handle this directly for the original
> > > > > guilty queue.
> > > > >  >>
> > > > >  >> Signed-off-by: Alex Deucher <[email protected]>  >> Cc:
> > > > > Srinivasan Shanmugam <[email protected]>  >> Cc: Prike
> > > > > Liang <[email protected]>  >> Cc: Sunil Khatri
> > > > > <[email protected]>  >> ---  >>
> > > > > drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 7 ++++++-  >>   1 file
> > > > > changed, 6 insertions(+), 1 deletion(-)  >>  >> diff --git
> > > > > a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> > > > >  >> b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> > > > >  >> index 6aa75da27f912..5e1262636e1e9 100644  >> ---
> > > > > a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> > > > >  >> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> > > > >  >> @@ -146,8 +146,13 @@ static void
> > > > > amdgpu_userq_hang_detect_work(struct
> > > > >  >> work_struct *work)
> > > > >  >>                                                         queue,
> > > > > NULL, NULL);  >>                else  >>                        r =
> > > > > userq_funcs->reset(queue);  >> -             if (r)  >> +
> > > > > if (r) {  >>                        gpu_reset = true;  >> +
> > > > > } else {  >> +
> > > > > atomic_inc(&adev->gpu_reset_counter);
> > > > >  >> +
> > > > > amdgpu_userq_fence_driver_force_completion(queue);
> > > > >  >> +                     drm_dev_wedged_event(adev_to_drm(adev),
> > > > >  >> DRM_WEDGE_RECOVERY_NONE, NULL);
> > > > >  >> +             }
> > > > >  >>        } else {
> > > > >  >>                gpu_reset = true;
> > > > >  >>        }
> > > > >  >
> > > > >  > After the original queue was reset successfully, it did not
> > > > > update gpu_reset_counter, complete its pending fences, or send the 
> > > > > wedged
> > event.
> > > > >  > mes_userq_reset_queue() already updates gpu_reset_counter,
> > > > > completes the pending fences, and sends the wedged event for the
> > > > > other affected queues,  > but skips the original queue because it
> > > > > has already been reset.
> > > > >
> > > > > What is the rationale of sending multiple device wedged events on
> > > > > a per queue basis?
> > > > >
> > > > > The question of whether drm_dev_wedged_event() should be emitted
> > > > > once per queue or once per overall recovery seems like a broader
> > > > > design discussion.
> > > > >
> > > >
> > > > Along with that, also need to consider if device reset_counter needs
> > > > to be incremented on a per queue basis or based on reset event
> > > > recovery. It could get incremented multiple times inside this -
> > mes_userq_reset_queue.
> > >
> > > Looking at the current flow, both gpu_reset_counter and
> > drm_dev_wedged_event() are updated once for each successfully reset queue. 
> > It
> > would be helpful to clarify whether they are intended to be updated per 
> > affected
> > queue or once per overall recovery.
> > >
> >
> > What are the semantics around the reset counter and wedged events?
> > Presumably each should be incremented for each queue that is reset? If a 
> > hang
> > affects multiple queues shouldn't each be a separate "reset"?
> > In the most common case, there should just be one since queue reset should 
> > be
> > able to reset just the guilty queue.
>
> Thanks for the clarification, Alex. Understood that the reset counter and 
> wedged event are intended to be updated once for each queue that is reset.

Well, I guess that is the question.  We are the semantics around these?

Alex

>
> Best Regards,
> Srini
>
> >
> > Alex

Reply via email to