On 2026-09-02 08:49, Prike Liang wrote: > The eviction fence suspend worker can submit MES REMOVE_QUEUE packets > without holding the reset-domain semaphore. If GPU recovery starts while > the worker is running, both paths can access the hardware concurrently. > This triggers the hardware-access lockdep assertion and can submit a MES > packet while recovery is resetting the device. > > Try to take the reset-domain semaphore for read around userq eviction. Do > not block on it while holding userq_mutex because recovery takes the reset > semaphore for write before acquiring buffer reservations and userq_mutex. > Instead, drop userq_mutex, wait for recovery without holding any other > lock, and retry the queue-state checks after recovery completes. > > This makes recovery wait for an in-flight MES eviction, while an eviction > which starts after recovery waits without introducing the reverse lock > dependency that caused the reported circular-lock warning.
It only changes amdgpu_eviction_fence_suspend_worker(). It adds reset_domain->sem around the evict path only. It does not touch amdgpu_userq_vm_validate_and_restore_queue() or amdgpu_evf_mgr_rearm(). Eviction is now serialized against reset, but restore/rearm is not. Same worker, still unprotected. Thanks, Vitaly > Signed-off-by: Prike Liang <[email protected]> > --- > drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c | 9 +++++++++ > 1 file changed, 9 insertions(+) > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c > b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c > index 2ea8553c82f0..93307cbf55dc 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c > @@ -67,10 +67,18 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct > *work) > bool cookie; > int r; > > +retry: > mutex_lock(&uq_mgr->userq_mutex); > > /* Fence waits are not allowed in a fence signalling critical section. > */ > amdgpu_userq_wait_for_signal(uq_mgr); > + if (!down_read_trylock(&uq_mgr->adev->reset_domain->sem)) { > + mutex_unlock(&uq_mgr->userq_mutex); > + > + down_read(&uq_mgr->adev->reset_domain->sem); > + up_read(&uq_mgr->adev->reset_domain->sem); > + goto retry; > + } > > /* > * This is intentionally after taking the userq_mutex since we do > @@ -81,6 +89,7 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct > *work) > > ev_fence = amdgpu_evf_mgr_get_fence(evf_mgr); > r = amdgpu_userq_evict(uq_mgr); > + up_read(&uq_mgr->adev->reset_domain->sem); > if (r) > dma_fence_set_error(ev_fence, r); >
