On 03-Sep-26 9:32 AM, Lazar, Lijo wrote:


On 02-Sep-26 10:25 PM, Deucher, Alexander wrote:
Public


For kernel queues each IB executes with a kernel provided vmid assigned dynamically by the kernel driver.


In this case, it's a device level TMA 'kq_tma_bo' for first level. That address is programmed in SQ registers. When a job is submitted, the second level handler is picked from what is programmed in kq_tma_bo. Do you mean to say that driver will change that value dynamically based on what is provided by user?

Do you mean to say that driver will change that value dynamically based on what is provided by user for each job submission?

Thanks,
Lijo


Thanks,
Lijo

  > Alex

*From:*Lazar, Lijo <[email protected]>
*Sent:* Wednesday, September 2, 2026 12:12 PM
*To:* SHANMUGAM, SRINIVASAN <[email protected]>; Koenig, Christian <[email protected]>; Deucher, Alexander <[email protected]>
*Cc:* [email protected]
*Subject:* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure

Public

I'm not sure how this works, I thought the TMA mapping is per VMID and kernel queues have static VMIDs.

Thanks,

Lijo

------------------------------------------------------------------------

*From:*SHANMUGAM, SRINIVASAN <[email protected] <mailto:[email protected]>>
*Sent:* Wednesday, 02 September 2026 21:29:38
*To:* Lazar, Lijo <[email protected] <mailto:[email protected]>>; Koenig, Christian <[email protected] <mailto:[email protected]>>; Deucher, Alexander <[email protected] <mailto:[email protected]>> *Cc:* [email protected] <mailto:amd- [email protected]> <[email protected] <mailto:amd- [email protected]>> *Subject:* RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure

Public

-----Original Message-----
From: Lazar, Lijo <[email protected] <mailto:[email protected]>>
Sent: Wednesday, September 2, 2026 8:59 PM
To: SHANMUGAM, SRINIVASAN <[email protected] <mailto:[email protected]>>; Koenig, Christian <[email protected] <mailto:[email protected]>>; Deucher,
Alexander
<[email protected] <mailto:[email protected]>>
Cc: [email protected] <mailto:[email protected]>
Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler
infrastructure



On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote:
> MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program
> SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the
> driver program the first-level CWSR trap handler for these VMIDs
> directly via SRBM select.
>
> Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for
> kernel queue VMIDs. Unlike per-process TMA (created in
> amdgpu_trap_alloc), this is device-level and lives for the lifetime of
> the device. It is zero-initialized by the OS: the second-level handler
> address is 0 until userspace calls SET_L2_TRAP.
>
> Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW
> vmhub callback, and a new program_kernel_trap_vmids hook in
> amdgpu_vmhub_funcs for per-GFX-generation register writes.
>
> Required for:
>    - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request)
>    - Consistent trap handler behavior when switching between kernel
>      queues and user queues
>
> Suggested-by: Christian König <[email protected] <mailto:[email protected]>> > Cc: Alexander Deucher <[email protected] <mailto:[email protected]>> > Signed-off-by: Srinivasan Shanmugam <[email protected] <mailto:[email protected]>>
> Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8
> ---
>   drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h  |  1 +
>   drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37
++++++++++++++++++++++++
>   drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h |  2 ++
>   3 files changed, 40 insertions(+)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> index 3ca187f5ade8..5624a5ab5c62 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs {
>     void (*print_l2_protection_fault_status)(struct amdgpu_device *adev,
>                                              uint32_t status);
>     uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t
> flush_type);
> +   void (*program_kernel_trap_vmids)(struct amdgpu_device *adev);
>   };
>
>   struct amdgpu_vmhub {
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> index 31f653ec3fb1..0e0aeea0aa2d 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev)
>
>     memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz);
>
> +   /*
> +    * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject > +    * to eviction. Zero-initialized by OS: second-level handler address
> +    * is 0 until userspace calls SET_L2_TRAP.

How is the exclusivity maintained as a user app doesn't 'own' kernel queue? How is the conflict of different user apps trying to install their own second level handler on a
kernel queue handled?

As pointed out by Alex:

   - Each process gets its own per-VM TMA buffer for kernel queues
   - It is allocated and mapped at a fixed VA in the process's GPUVM
     when the device is opened, similar to amdgpu_map_static_csa()
   - Before each job is dispatched to a kernel queue VMID, the driver
     programs SQ_SHADER_TMA to that process's own TMA VA

This way:
   - App A submits job → TMA = App A's TMA → job runs
   - App B submits job → TMA = App B's TMA → job runs
   - No conflict — each process has its own TMA buffer

Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series.
This series only installs the first-level trap handler.

For second-level handler support on kernel queues, I think since:

   - Each process has its own per-VM TMA buffer (allocated at device
     open, mapped at a fixed VA in the process's GPUVM)
   - User calls SET_L2_TRAP → writes second-level handler address
     into that process's own TMA buffer
   - When shader crashes → first-level handler reads from that
     process's TMA → jumps to that process's second-level handler
   - No conflict — each process has its own TMA with its own
     second-level handler address

The per-VM TMA design is the foundation for this future series.
The device-level kq_tma_bo will be removed in v2.

For long term — kernel queues are replaced by user queues entirely
, where MES already handles this correctly via ADD_QUEUE.

Thanks,
Srini



Reply via email to