When amdgpu_vm_ptes_update fails, there might be an IB jobs leak.

This was built on top of the patchset at [1].

A first attempt at fixing this was submitted at [2]. In that thread, it
was discussed that the failure happens due to an unmap trying to split a
hugepage and trying to allocate a PT BO. This is fixed by the patchset
at [1], by making the unmap not allocate any BOs.

However, the same failure might happen later, when the new mappings are
setup and the PT BO needs to be allocated.

The approach taken here is that instead of doing a commit, we introduce
a new abort operation and call it. In the case of SDMA, it frees the
job.

The commit still refers to amdgpu_vm_update_range, but the leak has been
reproduced with the patchset, so it shows up at amdgpu_vm_map_range. I
kept the original function name in the commit, however, as we might
consider to backport this.

[1] 
https://lore.kernel.org/amd-gfx/[email protected]/
[2] 
https://lore.kernel.org/amd-gfx/[email protected]/

Signed-off-by: Thadeu Lima de Souza Cascardo <[email protected]>
---
Thadeu Lima de Souza Cascardo (2):
      drm/amdgpu: make amdgpu_vm_update_funcs commit return void
      drm/amdgpu: free job when amdgpu_vm_ptes_update fails

 drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c          | 24 ++++++++++++------------
 drivers/gpu/drm/amd/amdgpu/amdgpu_vm_cpu.c      | 13 ++++++++-----
 drivers/gpu/drm/amd/amdgpu/amdgpu_vm_internal.h |  5 +++--
 drivers/gpu/drm/amd/amdgpu/amdgpu_vm_pt.c       |  9 ++++++---
 drivers/gpu/drm/amd/amdgpu/amdgpu_vm_sdma.c     | 20 +++++++++++---------
 5 files changed, 40 insertions(+), 31 deletions(-)
---
base-commit: a7629c76ab981c0d5749eef9f4e47d12a2083330
change-id: 20260929-amdgpu_vm_update_funcs_abort-97cb4349b68d

Best regards,
--  
Thadeu Lima de Souza Cascardo <[email protected]>

Reply via email to