On Tue, Jul 14, 2026 at 10:10 AM SHANMUGAM, SRINIVASAN <[email protected]> wrote: > > AMD General > > > -----Original Message----- > > From: Alex Deucher <[email protected]> > > Sent: Tuesday, July 14, 2026 7:26 PM > > To: SHANMUGAM, SRINIVASAN <[email protected]> > > Cc: Lazar, Lijo <[email protected]>; Deucher, Alexander > > <[email protected]>; [email protected]; Liang, Prike > > <[email protected]>; Khatri, Sunil <[email protected]> > > Subject: Re: [PATCH] drm/amdgpu/userq: properly account for resets > > > > On Tue, Jul 14, 2026 at 9:49 AM SHANMUGAM, SRINIVASAN > > <[email protected]> wrote: > > > > > > AMD General > > > > > > > -----Original Message----- > > > > From: Lazar, Lijo <[email protected]> > > > > Sent: Tuesday, July 14, 2026 4:02 PM > > > > To: SHANMUGAM, SRINIVASAN <[email protected]>; > > Deucher, > > > > Alexander <[email protected]>; amd- > > > > [email protected] > > > > Cc: Liang, Prike <[email protected]>; Khatri, Sunil > > > > <[email protected]> > > > > Subject: Re: [PATCH] drm/amdgpu/userq: properly account for resets > > > > > > > > > > > > > > > > On 14-Jul-26 3:57 PM, SHANMUGAM, SRINIVASAN wrote: > > > > > AMD General > > > > > > > > > > > > > > > > > > > > > > > > > Get Outlook for Android <https://aka.ms/AAb9ysg> > > > > > > > > > > ------------------------------------------------------------------ > > > > > ---- > > > > > -- > > > > > *From:* Lazar, Lijo <[email protected]> > > > > > *Sent:* Tuesday, July 14, 2026 3:14:34 PM > > > > > *To:* SHANMUGAM, SRINIVASAN > > <[email protected]>; > > > > Deucher, > > > > > Alexander <[email protected]>; > > > > > [email protected] <[email protected]> > > > > > *Cc:* Liang, Prike <[email protected]>; Khatri, Sunil > > > > > <[email protected]> > > > > > *Subject:* Re: [PATCH] drm/amdgpu/userq: properly account for > > > > > resets > > > > > > > > > > > > > > > > > > > > On 14-Jul-26 10:16 AM, SHANMUGAM, SRINIVASAN wrote: > > > > > > AMD General > > > > > > > > > > > >> -----Original Message----- > > > > > >> From: Alex Deucher <[email protected]> >> Sent: > > > > > Tuesday, July 14, 2026 2:09 AM >> To: [email protected] > > > > > >> > > Cc: > > > > > Deucher, Alexander <[email protected]>; SHANMUGAM, >> > > > > > SRINIVASAN <[email protected]>; Liang, Prike >> > > > > > <[email protected]>; Khatri, Sunil <[email protected]> >> > > > > > Subject: [PATCH] drm/amdgpu/userq: properly account for resets >> > > > > > >> We need to increment the reset counter, force fence completion, > > > > > and set the >> wedged event when a user queue is reset. > > > > > >> > > > > > >> mes_userq_reset_queue() handles this for collateral damage, > > > > > but the caller needs >> to handle this directly for the original > > > > > guilty queue. > > > > > >> > > > > > >> Signed-off-by: Alex Deucher <[email protected]> >> Cc: > > > > > Srinivasan Shanmugam <[email protected]> >> Cc: Prike > > > > > Liang <[email protected]> >> Cc: Sunil Khatri > > > > > <[email protected]> >> --- >> > > > > > drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 7 ++++++- >> 1 file > > > > > changed, 6 insertions(+), 1 deletion(-) >> >> diff --git > > > > > a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c > > > > > >> b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c > > > > > >> index 6aa75da27f912..5e1262636e1e9 100644 >> --- > > > > > a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c > > > > > >> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c > > > > > >> @@ -146,8 +146,13 @@ static void > > > > > amdgpu_userq_hang_detect_work(struct > > > > > >> work_struct *work) > > > > > >> queue, > > > > > NULL, NULL); >> else >> r = > > > > > userq_funcs->reset(queue); >> - if (r) >> + > > > > > if (r) { >> gpu_reset = true; >> + > > > > > } else { >> + > > > > > atomic_inc(&adev->gpu_reset_counter); > > > > > >> + > > > > > amdgpu_userq_fence_driver_force_completion(queue); > > > > > >> + drm_dev_wedged_event(adev_to_drm(adev), > > > > > >> DRM_WEDGE_RECOVERY_NONE, NULL); > > > > > >> + } > > > > > >> } else { > > > > > >> gpu_reset = true; > > > > > >> } > > > > > > > > > > > > After the original queue was reset successfully, it did not > > > > > update gpu_reset_counter, complete its pending fences, or send the > > > > > wedged > > event. > > > > > > mes_userq_reset_queue() already updates gpu_reset_counter, > > > > > completes the pending fences, and sends the wedged event for the > > > > > other affected queues, > but skips the original queue because it > > > > > has already been reset. > > > > > > > > > > What is the rationale of sending multiple device wedged events on > > > > > a per queue basis? > > > > > > > > > > The question of whether drm_dev_wedged_event() should be emitted > > > > > once per queue or once per overall recovery seems like a broader > > > > > design discussion. > > > > > > > > > > > > > Along with that, also need to consider if device reset_counter needs > > > > to be incremented on a per queue basis or based on reset event > > > > recovery. It could get incremented multiple times inside this - > > mes_userq_reset_queue. > > > > > > Looking at the current flow, both gpu_reset_counter and > > drm_dev_wedged_event() are updated once for each successfully reset queue. > > It > > would be helpful to clarify whether they are intended to be updated per > > affected > > queue or once per overall recovery. > > > > > > > What are the semantics around the reset counter and wedged events? > > Presumably each should be incremented for each queue that is reset? If a > > hang > > affects multiple queues shouldn't each be a separate "reset"? > > In the most common case, there should just be one since queue reset should > > be > > able to reset just the guilty queue. > > Thanks for the clarification, Alex. Understood that the reset counter and > wedged event are intended to be updated once for each queue that is reset.
Well, I guess that is the question. We are the semantics around these? Alex > > Best Regards, > Srini > > > > > Alex
