On 2026-09-02 08:49, Prike Liang wrote:
> The eviction fence suspend worker can submit MES REMOVE_QUEUE packets
> without holding the reset-domain semaphore. If GPU recovery starts while
> the worker is running, both paths can access the hardware concurrently.
> This triggers the hardware-access lockdep assertion and can submit a MES
> packet while recovery is resetting the device.
>
> Try to take the reset-domain semaphore for read around userq eviction. Do
> not block on it while holding userq_mutex because recovery takes the reset
> semaphore for write before acquiring buffer reservations and userq_mutex.
> Instead, drop userq_mutex, wait for recovery without holding any other
> lock, and retry the queue-state checks after recovery completes.
>
> This makes recovery wait for an in-flight MES eviction, while an eviction
> which starts after recovery waits without introducing the reverse lock
> dependency that caused the reported circular-lock warning.

It only changes amdgpu_eviction_fence_suspend_worker(). It adds
reset_domain->sem around the evict path only.

It does not touch amdgpu_userq_vm_validate_and_restore_queue() or
amdgpu_evf_mgr_rearm().

Eviction is now serialized against reset, but restore/rearm
is not. Same worker, still unprotected.

Thanks, Vitaly

> Signed-off-by: Prike Liang <[email protected]>
> ---
>  drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c | 9 +++++++++
>  1 file changed, 9 insertions(+)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c 
> b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
> index 2ea8553c82f0..93307cbf55dc 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
> @@ -67,10 +67,18 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct 
> *work)
>       bool cookie;
>       int r;
>  
> +retry:
>       mutex_lock(&uq_mgr->userq_mutex);
>  
>       /* Fence waits are not allowed in a fence signalling critical section. 
> */
>       amdgpu_userq_wait_for_signal(uq_mgr);
> +     if (!down_read_trylock(&uq_mgr->adev->reset_domain->sem)) {
> +             mutex_unlock(&uq_mgr->userq_mutex);
> +
> +             down_read(&uq_mgr->adev->reset_domain->sem);
> +             up_read(&uq_mgr->adev->reset_domain->sem);
> +             goto retry;
> +     }
>  
>       /*
>        * This is intentionally after taking the userq_mutex since we do
> @@ -81,6 +89,7 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct 
> *work)
>  
>       ev_fence = amdgpu_evf_mgr_get_fence(evf_mgr);
>       r = amdgpu_userq_evict(uq_mgr);
> +     up_read(&uq_mgr->adev->reset_domain->sem);
>       if (r)
>               dma_fence_set_error(ev_fence, r);
>  

Reply via email to