Otherwise, we might leak the SDMA job.

[ 2877.877264] kmemleak: unreferenced object 0xffff89b1f3cf0800 (size 1024):
[ 2877.877273] kmemleak:   comm "vm_always_valid", pid 3027, jiffies 4295714403
[ 2877.877275] kmemleak:   hex dump (first 32 bytes):
[ 2877.877277] kmemleak:     00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00  
................
[ 2877.877279] kmemleak:     00 03 da 52 b1 89 ff ff 38 25 5f cb b1 89 ff ff  
...R....8%_.....
[ 2877.877280] kmemleak:   backtrace (crc d52b9041):
[ 2877.877282] kmemleak:     __kmalloc_noprof+0x4b3/0x770
[ 2877.877288] kmemleak:     amdgpu_job_alloc+0x69/0x280 [amdgpu]
[ 2877.877668] kmemleak:     amdgpu_job_alloc_with_ib+0x55/0xf0 [amdgpu]
[ 2877.878023] kmemleak:     amdgpu_vm_sdma_prepare+0x4b/0xa0 [amdgpu]
[ 2877.878350] kmemleak:     amdgpu_vm_update_range+0x24f/0x940 [amdgpu]
[ 2877.878668] kmemleak:     amdgpu_vm_clear_freed+0x13b/0x290 [amdgpu]
[ 2877.878983] kmemleak:     amdgpu_gem_object_close+0x1ac/0x270 [amdgpu]
[ 2877.879300] kmemleak:     drm_gem_object_release_handle+0x35/0xd0
[ 2877.879305] kmemleak:     idr_for_each+0x70/0xe0
[ 2877.879310] kmemleak:     drm_gem_release+0x23/0x30
[ 2877.879311] kmemleak:     drm_file_free+0x217/0x2a0
[ 2877.879314] kmemleak:     drm_release+0x61/0xe0
[ 2877.879317] kmemleak:     amdgpu_drm_release+0x62/0xd0 [amdgpu]
[ 2877.879630] kmemleak:     __fput+0xfb/0x2d0
[ 2877.879634] kmemleak:     __x64_sys_close+0x3d/0x80
[ 2877.879637] kmemleak:     do_syscall_64+0x12d/0x6c0

Also, flush the TLB and free the entries in the local stack
tlb_flush_waitlist. Otherwise we might touch old stack addresses when
removing the entries when we finish the VM.

[ 2791.162517] ------------[ cut here ]------------
[ 2791.162519] list_del corruption. next->prev should be ffff89b27cf0b778, but 
was 0000000000000000. (next=ffffcacdc12e77c8)
[ 2791.162522] WARNING: lib/list_debug.c:65 at 
__list_del_entry_valid_or_report+0xea/0x100, CPU#1: vm_always_valid/3027
[ 2791.162659] CPU: 1 UID: 1000 PID: 3027 Comm: vm_always_valid Tainted: G      
  W           7.2.0-02902-g432c2ced2942 #11 PREEMPT  
2356ac135df40d3c9080402059c7492776a63ecd
[ 2791.162663] Tainted: [W]=WARN
[ 2791.162665] Hardware name: Valve Jupiter/Jupiter, BIOS F7A0133 08/05/2024
[ 2791.162668] RIP: 0010:__list_del_entry_valid_or_report+0xf4/0x100
[...]
[ 2791.162688] Call Trace:
[ 2791.162690]  <TASK>
[ 2791.162694]  amdgpu_vm_pt_free+0x5d/0xa0 [amdgpu 
09338d95446a22cfd002ef2392fde61d2ef7b139]
[ 2791.163070]  amdgpu_vm_pt_free_root+0xe3/0x130 [amdgpu 
09338d95446a22cfd002ef2392fde61d2ef7b139]
[ 2791.163392]  amdgpu_vm_fini+0x320/0x5f0 [amdgpu 
09338d95446a22cfd002ef2392fde61d2ef7b139]
[ 2791.163709]  ? __xa_erase+0x57/0xa0
[ 2791.163714]  ? _raw_spin_unlock_irqrestore+0x34/0x60
[ 2791.163718]  ? _raw_spin_unlock_irqrestore+0x34/0x60
[ 2791.163720]  ? trace_hardirqs_on+0x16/0xd0
[ 2791.163725]  ? _raw_spin_unlock_irqrestore+0x3f/0x60
[ 2791.163728]  amdgpu_driver_postclose_kms+0x1d1/0x2d0 [amdgpu 
09338d95446a22cfd002ef2392fde61d2ef7b139]
[ 2791.164042]  drm_file_free+0x23a/0x2a0
[ 2791.164049]  drm_release+0x61/0xe0
[ 2791.164052]  amdgpu_drm_release+0x62/0xd0 [amdgpu 
09338d95446a22cfd002ef2392fde61d2ef7b139]
[ 2791.164362]  __fput+0xfb/0x2d0
[ 2791.164367]  __x64_sys_close+0x3d/0x80
[ 2791.164371]  do_syscall_64+0x12d/0x6c0

This can be caused by VRAM memory pressure, where a PT BO allocation
fails during amdgpu_gem_va_ioctl.

Since the failure from amdgpu_vm_update_range is still returned,
amdgpu_vm_bo_update will leave the mappings at the invalids list and
amdgpu_vm_clear_freed will keep them at the freed list. This will allow
amdgpu_cs_ioctl to retry later.

Fixes: 81417bea8755 ("drm/amdgpu: explicitly sync VM update to PDs/PTs")
Fixes: c3546695830e ("drm/amdgpu: use the new VM backend for PTEs")
Fixes: b6c4f90b3819 ("drm/amdgpu: sync page table freeing with tlb flush")
Signed-off-by: Thadeu Lima de Souza Cascardo <[email protected]>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c 
b/drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c
index 5bde36754607..1495adaaacd8 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c
@@ -1235,7 +1235,7 @@ int amdgpu_vm_update_range(struct amdgpu_device *adev, 
struct amdgpu_vm *vm,
                tmp = start + num_entries;
                r = amdgpu_vm_ptes_update(&params, start, tmp, addr, flags);
                if (r)
-                       goto error_free;
+                       break;
 
                amdgpu_res_next(&cursor, num_entries * AMDGPU_GPU_PAGE_SIZE);
                start = tmp;

-- 
2.47.3

Reply via email to