MES (Micro Engine Scheduler) does not perform heavy-weight TLB
invalidation after unmapping queues, unlike HWS which does this
automatically. This causes a race condition where in-flight DMA
descriptors can access memory that has been unmapped, leading to page
faults and GPU queue hangs during SVM page migration.

The issue manifests as KFDSVMRangeTest.MultiThreadMigrationTest
failures on gfx1151 (Strix Point) with XNACK mode 1 enabled - the GPU
compute queue hangs with packets submitted but never consumed.

Add kfd_flush_tlb() calls after MES queue removal in two locations:
- evict_process_queues_cpsch(): after all queues removed during eviction
- suspend_queues(): after debug/criu queue suspension (with mem_fence barrier)

This ensures all in-flight memory accesses from unmapped queues are
flushed before memory is freed or migrated.

Change-Id: Iaa884d4a7ad1199446bc45f4ad8a9a179ab386e6
Signed-off-by: Priya Hosur <[email protected]>
Reviewed-by: Felix Kuehling <[email protected]>
---
 .../gpu/drm/amd/amdkfd/kfd_device_queue_manager.c   | 13 ++++++++++++-
 1 file changed, 12 insertions(+), 1 deletion(-)

diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c 
b/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
index a23384571193..6002c8a65fbe 100644
--- a/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
+++ b/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
@@ -1450,6 +1450,14 @@ static int evict_process_queues_cpsch(struct 
device_queue_manager *dqm,
                dqm_evict_mqd_bo(dqm, q);
        }
 
+       /*
+        * Heavy-weight TLB flush after MES removes queues to ensure
+        * in-flight memory accesses complete before memory is freed/migrated.
+        * HWS does this automatically, MES does not.
+        */
+       if (dqm->dev->kfd->shared_resources.enable_mes)
+               kfd_flush_tlb(pdd);
+
        if (!dqm->dev->kfd->shared_resources.enable_mes) {
                pdd->last_evict_timestamp = get_jiffies_64();
                retval = execute_queues_cpsch(dqm,
@@ -3736,8 +3744,11 @@ int suspend_queues(struct kfd_process *p,
                if (!per_device_suspended) {
                        dqm_unlock(dqm);
                        mutex_unlock(&p->event_mutex);
-                       if (total_suspended)
+                       if (total_suspended) {
                                amdgpu_amdkfd_debug_mem_fence(dqm->dev->adev);
+                               /* Heavy-weight TLB flush after MES suspends 
queues */
+                               kfd_flush_tlb(pdd);
+                       }
                        continue;
                }
 
-- 
2.43.0

Reply via email to