Public > -----Original Message----- > From: amd-gfx <[email protected]> On Behalf Of > Srinivasan Shanmugam > Sent: Wednesday, September 2, 2026 11:07 AM > To: Koenig, Christian <[email protected]>; Deucher, Alexander > <[email protected]> > Cc: [email protected]; SHANMUGAM, SRINIVASAN > <[email protected]> > Subject: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler > infrastructure > > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the driver > program the first-level CWSR trap handler for these VMIDs directly via SRBM > select. > > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for > kernel queue VMIDs. Unlike per-process TMA (created in amdgpu_trap_alloc), > this is device-level and lives for the lifetime of the device. It is > zero-initialized by > the OS: the second-level handler address is 0 until userspace calls > SET_L2_TRAP.
There should be one BO per user GPUVM instance, and it needs to be mapped at the same address in every user's GPUVM address space. E.g., this should be allocated and mapped when the user's GPUVM object is created. Alex > > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW > vmhub callback, and a new program_kernel_trap_vmids hook in > amdgpu_vmhub_funcs for per-GFX-generation register writes. > > Required for: > - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request) > - Consistent trap handler behavior when switching between kernel > queues and user queues > > Suggested-by: Christian König <[email protected]> > Cc: Alexander Deucher <[email protected]> > Signed-off-by: Srinivasan Shanmugam <[email protected]> > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8 > --- > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 + > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37 > ++++++++++++++++++++++++ > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++ > 3 files changed, 40 insertions(+) > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > index 3ca187f5ade8..5624a5ab5c62 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs { > void (*print_l2_protection_fault_status)(struct amdgpu_device *adev, > uint32_t status); > uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t > flush_type); > + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev); > }; > > struct amdgpu_vmhub { > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > index 31f653ec3fb1..0e0aeea0aa2d 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev) > > memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz); > > + /* > + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not > subject > + * to eviction. Zero-initialized by OS: second-level handler address > + * is 0 until userspace calls SET_L2_TRAP. > + */ > + r = amdgpu_bo_create_kernel(adev, > AMDGPU_TRAP_TMA_MAX_SIZE, PAGE_SIZE, > + AMDGPU_GEM_DOMAIN_GTT, > + &trap_info->kq_tma_bo, NULL, NULL); > + if (r) { > + /* isa_bo freed explicitly; trap_info struct freed by __free */ > + amdgpu_bo_free_kernel(&trap_info->isa_bo, NULL, NULL); > + return r; > + } > + > amdgpu_trap_cwsr_init_save_area_info(adev, trap_info); > adev->trap_info = no_free_ptr(trap_info); > + amdgpu_trap_program_kernel_vmids(adev); > > return 0; > } > @@ -267,11 +282,33 @@ void amdgpu_trap_fini(struct amdgpu_device > *adev) > if (!amdgpu_trap_is_enabled(adev)) > return; > > + amdgpu_bo_free_kernel(&adev->trap_info->kq_tma_bo, NULL, > NULL); > amdgpu_bo_free_kernel(&adev->trap_info->isa_bo, NULL, NULL); > kfree(adev->trap_info); > adev->trap_info = NULL; > } > > +/** > + * amdgpu_trap_program_kernel_vmids - program first-level trap handler for > + * kernel queue VMIDs > + * @adev: amdgpu device pointer > + * > + * Programs SQ_SHADER_TBA/TMA for kernel queue VMIDs > +(1..first_kfd_vmid-1) > + * via SRBM select. MES owns these VMIDs but does not program trap > +handler > + * state. Called after trap init and on GPU resume via setup_vmid_config. > + */ > +void amdgpu_trap_program_kernel_vmids(struct amdgpu_device *adev) { > + struct amdgpu_vmhub *hub = &adev- > >vmhub[AMDGPU_GFXHUB(0)]; > + > + if (!amdgpu_trap_is_enabled(adev)) > + return; > + if (!hub->vmhub_funcs || !hub->vmhub_funcs- > >program_kernel_trap_vmids) > + return; > + > + hub->vmhub_funcs->program_kernel_trap_vmids(adev); > +} > + > /* > * amdgpu_map_cwsr_trap_handler should be called during amdgpu_vm_init > * it maps virtual address amdgpu_trap_tba_vaddr() to this VM, and each diff > --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h > index 6d4664469bad..326910f1d94d 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h > @@ -48,6 +48,7 @@ struct amdgpu_trap_obj { struct amdgpu_trap_info { > /* cwsr isa */ > struct amdgpu_bo *isa_bo; > + struct amdgpu_bo *kq_tma_bo; /* pinned GTT, device-level > TMA for kernel queue VMIDs */ > const void *isa_buf; > uint32_t isa_sz; > /* cwsr size info per XCC*/ > @@ -70,6 +71,7 @@ struct amdgpu_trap_usr_addr { }; > > int amdgpu_trap_init(struct amdgpu_device *adev); > +void amdgpu_trap_program_kernel_vmids(struct amdgpu_device *adev); > void amdgpu_trap_fini(struct amdgpu_device *adev); > > int amdgpu_trap_alloc(struct amdgpu_device *adev, struct amdgpu_vm *vm, > -- > 2.34.1
