Like the MES scheduler ring, the KIQ ring sets no_scheduler = true and uses a
polling fence, so it is skipped by the force-completion loop in
amdgpu_device_pre_asic_reset(). Its hw fence value lives in wb (GTT) memory and
survives a MODE1 reset while fence_drv.sync_seq keeps advancing, so after a
reset the first KIQ submission can poll forever on a seq that is never written
back.

Force complete the KIQ ring fences too so their hw fence is realigned to
sync_seq.

Suggested-by: Alex Deucher <[email protected]>
Signed-off-by: Jesse Zhang <[email protected]>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_device.c | 12 ++++++++++++
 1 file changed, 12 insertions(+)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c 
b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
index 168947747c5c..77426e814e08 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
@@ -5049,6 +5049,18 @@ int amdgpu_device_pre_asic_reset(struct amdgpu_device 
*adev,
                        amdgpu_fence_driver_force_completion(mes_ring, fence);
        }
 
+       /*
+        * KIQ rings are polling-fence/no_scheduler like MES, so realign their
+        * fence too (one ring per XCC), otherwise the first post-reset KIQ
+        * submission polls forever on a stale seq.
+        */
+       for (i = 0; i < AMDGPU_MAX_GC_INSTANCES; i++) {
+               struct amdgpu_ring *kiq_ring = &adev->gfx.kiq[i].ring;
+
+               if (kiq_ring->fence_drv.initialized && kiq_ring->sched.ready)
+                       amdgpu_fence_driver_force_completion(kiq_ring, fence);
+       }
+
        amdgpu_fence_driver_isr_toggle(adev, false);
 
        r = amdgpu_reset_prepare_hwcontext(adev, reset_context);
-- 
2.49.0

Reply via email to