On 2026-08-12 02:11, Priya Hosur wrote:
MES (Micro Engine Scheduler) does not perform heavy-weight TLB invalidation
after unmapping queues, unlike HWS which does this automatically. This causes
a race condition where in-flight SDMA DMA descriptors can access memory that
has been unmapped, leading to page faults and GPU queue hangs during SVM
page migration.

The issue manifests as KFDSVMRangeTest.MultiThreadMigrationTest/1 failures
on gfx1151 (Ryzen AI MAX) with XNACK mode 1 enabled - the GPU compute queue
hangs with packets submitted but never consumed.

Add kfd_flush_tlb() with TLB_FLUSH_HEAVYWEIGHT in two MES code paths:
1. evict_process_queues_cpsch() - after MES removes queues
2. suspend_queues() - after MES suspends queues and mem_fence completes

This ensures all in-flight memory accesses from unmapped queues are flushed
before memory is freed or migrated.

Testing on gfx1151 shows this reduces failure rate from 100% to approximately
7-10%. The residual failures require further investigation.

Signed-off-by: Priya Hosur <[email protected]>
---
  .../gpu/drm/amd/amdkfd/kfd_device_queue_manager.c   | 13 ++++++++++++-
  1 file changed, 12 insertions(+), 1 deletion(-)

diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c 
b/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
index a23384571193..5eb85290126e 100644
--- a/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
+++ b/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
@@ -1450,6 +1450,14 @@ static int evict_process_queues_cpsch(struct 
device_queue_manager *dqm,
                dqm_evict_mqd_bo(dqm, q);
        }
+ /*
+        * Heavy-weight TLB flush after MES removes queues to ensure
+        * in-flight SDMA accesses complete before memory is freed/migrated.
+        * HWS does this automatically, MES does not.

I'm not sure why you call out SDMA specifically here. This affects in-flight memory accesses from compute jobs as well. Just remove "SDMA" from the comment. With that fixed, the patch is

Reviewed-by: Felix Kuehling <[email protected]>


+        */
+       if (dqm->dev->kfd->shared_resources.enable_mes)
+               kfd_flush_tlb(pdd, TLB_FLUSH_HEAVYWEIGHT);
+
        if (!dqm->dev->kfd->shared_resources.enable_mes) {
                pdd->last_evict_timestamp = get_jiffies_64();
                retval = execute_queues_cpsch(dqm,
@@ -3736,8 +3744,11 @@ int suspend_queues(struct kfd_process *p,
                if (!per_device_suspended) {
                        dqm_unlock(dqm);
                        mutex_unlock(&p->event_mutex);
-                       if (total_suspended)
+                       if (total_suspended) {
                                amdgpu_amdkfd_debug_mem_fence(dqm->dev->adev);
+                               /* Heavy-weight TLB flush after MES suspends 
queues */
+                               kfd_flush_tlb(pdd, TLB_FLUSH_HEAVYWEIGHT);
+                       }
                        continue;
                }

Reply via email to