On Wed, Jul 8, 2026 at 10:54 AM Lazar, Lijo <[email protected]> wrote: > > > > On 08-Jul-26 7:29 PM, Deucher, Alexander wrote: > > AMD General > > > >> -----Original Message----- > >> From: Lazar, Lijo <[email protected]> > >> Sent: Tuesday, July 7, 2026 11:34 AM > >> To: Deucher, Alexander <[email protected]>; amd- > >> [email protected] > >> Subject: Re: [PATCH 4/6] drm/amdgpu: disallow gfxoff around > >> gfx/compute/sdma KQ reset > >> > >> > >> > >> On 07-Jul-26 8:04 PM, Alex Deucher wrote: > >>> If the engine is hung, we shouldn't need this since it won't be > >>> powered down, but we could potentially have a case where a queue can > >>> be preempted, but due to logic errors in the commands, the fence never > >>> signals. The engine could potentially go into the off state. Protect > >>> against that. > >> > >> On a second thought, is this a real case? If engine already goes to off > >> state, > >> wouldn't everything be reset by default once it comes back? > > > > I'm not sure if gfxoff will actually enter or not, but consider the case of > > an IB which is waiting for some memory location. It's just waiting so the > > queue can be preempted, but when it's resumed, it just goes back to waiting > > so it needs to be reset even though it's not hung per se because it's fence > > will never signal. For compute queues at least, we peempt all the queues > > in the reset handler and then reset the bad queues. In theory there could > > be a window there where gfxoff entry could happen and then we may access a > > GC register via MMIO when we reset one of the bad queues. > > > > I'm assuming IB execution will be considered as some part of GFX being > active. A wakeup from gfxoff requires external intervention, so it's > unlikely that it enters gfxoff during the wait. As far as I know, only > doorbells could trigger external wakeup.
Couldn't gfxoff entry happen after MES preempts queues? if we ring the MES doorbell subsequently it should cause gfxoff entry, but in some cases, we do the reset via MMIO. Alex > > Thanks, > Lijo > > Alex > > > >> > >> Thanks, > >> Lijo > >> > >>> > >>> Signed-off-by: Alex Deucher <[email protected]> > >>> --- > >>> drivers/gpu/drm/amd/amdgpu/amdgpu_job.c | 8 ++++++++ > >>> 1 file changed, 8 insertions(+) > >>> > >>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c > >>> b/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c > >>> index b5c2d1d4d8839..c8a1aa88a9de0 100644 > >>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c > >>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c > >>> @@ -137,7 +137,15 @@ static enum drm_gpu_sched_stat > >> amdgpu_job_timedout(struct drm_sched_job *s_job) > >>> s_job->sched->name); > >>> /* Stop the scheduler to prevent anybody else from touching > >> the ring buffer. */ > >>> drm_sched_wqueue_stop(&ring->sched); > >>> + if (ring->funcs->type == AMDGPU_RING_TYPE_GFX || > >>> + ring->funcs->type == AMDGPU_RING_TYPE_COMPUTE || > >>> + ring->funcs->type == AMDGPU_RING_TYPE_SDMA) > >>> + amdgpu_gfx_off_ctrl(adev, false); > >>> r = amdgpu_ring_reset(ring, job->vmid, job->hw_fence); > >>> + if (ring->funcs->type == AMDGPU_RING_TYPE_GFX || > >>> + ring->funcs->type == AMDGPU_RING_TYPE_COMPUTE || > >>> + ring->funcs->type == AMDGPU_RING_TYPE_SDMA) > >>> + amdgpu_gfx_off_ctrl(adev, true); > >>> if (!r) { > >>> /* Start the scheduler again */ > >>> drm_sched_wqueue_start(&ring->sched); > > >
