Hi, Yes, the signature matches. ttm_resource.c:235 in 7.2.y is the WARN_ON(pos->first->bo->base.resv != res->bo->base.resv) in ttm_lru_bulk_move_add(): pos->first points at a freed resource, the resv comparison reads garbage, and ttm_lru_bulk_move_pos_tail() then dereferences the stale pos->last — hence the NULL deref at +0xc9. The oops happens under bdev->lru_lock, which is why the KFD worker and amdgpu_drm_release then hang on spinlocks.
Hibernation was just my trigger; ttm_device_prepare_hibernation() is a global swapout, same ttm_bo_swapout_cb() path as ttm_global_swapout() under memory pressure. So yes, any swapout pass can leave the dangling endpoint — your report is the first hit I know of on a dGPU and without suspend. If you can, please add the oops to drm/amd#5387 <https://github.com/drm/amd/issues/5387>, it's useful data. On 7.2.y: the backport is in the 7.2.8-rc1 review series as "[PATCH 7.2 436/438] drm/ttm: fix swapped-out resources never leaving their bulk_move range" (review closes 25 Sep 14:05 UTC, so 7.2.8 should follow shortly). 7.2.y has no ttm_resource_try_charge(), so 3db7d7d can't be cherry-picked there at all (Greg's bot reported FAILED); the stable patch is the two commits squashed into the single swapout-site hunk. Your point (b) still applies to anyone hand-picking from mainline — a tree carrying only 3db7d7d is unfixed. Thanks for the detailed analysis. Vadim пт, 25 сент. 2026 г. в 12:03, 赤岡悠 <[email protected]>: > Hi, > > I hit a kernel Oops in the AMDGPU/TTM swapout path on a Radeon RX 7800 XT. > The > failure happened immediately after a global OOM condition and was followed > by an > RCU stall and a host reset. > > The kernel was `7.2.4-1-cachyos` on CachyOS. The GPU is a Sapphire Navi32 > / RX > 7800 XT (`1002:747e`, subsystem `1da2:475d`). The relevant userspace > versions > were Mesa `26.2.2-2`, KWin/Plasma `6.7.5-1.1`, and libdrm `2.4.134-1.1`. > > The first failure in the affected boot was: > > kswapd0 invoked oom-killer: gfp_mask=0xcc0(GFP_KERNEL), order=0 > oom-kill: ... global_oom ... task_memcg=<redacted>, task=llama-server > Out of memory: Killed process <pid> (llama-server) > > amdgpu 0000:03:00.0: [drm] *ERROR* Not enough memory for command > submission! > WARNING: drivers/gpu/drm/ttm/ttm_resource.c:235 at > ttm_resource_add_bulk_move+0x124/0x150 [ttm], CPU#15: QSGRenderThread > /2457 > BUG: kernel NULL pointer dereference, address: 0000000000000008 > Oops: Oops: 0000 [#1] SMP NOPTI > RIP: ttm_resource_add_bulk_move+0xc9/0x150 [ttm] > > The call chain was: > > ttm_resource_add_bulk_move > ttm_resource_alloc > ttm_bo_swapout_cb > ttm_lru_walk_for_evict > ttm_bo_swapout > ttm_global_swapout > ttm_tt_populate > ttm_bo_populate > ttm_bo_vm_fault_reserved > amdgpu_gem_fault > > After the Oops, an `amdgpu` KFD cleanup worker and an `amdgpu_drm_release` > path were both blocked on TTM-related spinlocks. The journal does not > contain > `Kernel panic - not syncing`; I am calling this an Oops followed by a > lockup or > reset, not a confirmed formal panic. > > This is consistent with the stale `bulk_move` endpoint bug introduced by > `b2ed01e7ad3d` ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout > failure", v7.1+): `ttm_tt_swapout()` returns the number of pages swapped on > success (never zero for a populated ttm), so the > `ttm_resource_del_bulk_move_unevictable()` / > `ttm_resource_move_to_lru_tail()` pair under `if (!ret)` is skipped on > every > successful swapout. A swapped-out resource stays inside its BO's bulk_move > range while unevictable; a later `ttm_resource_free()` / > `ttm_bo_set_bulk_move()` skips it via the `!ttm_resource_unevictable()` > guard, leaving a dangling range endpoint. Upstream names > `ttm_resource_add_bulk_move()` as one of the places that trips over that > cursor (a use-after-free on the dangling endpoint) — exactly the function > my NULL dereference hit at `+0xc9`. > > The trigger on my system differs from the hibernation cycles reported in > https://gitlab.freedesktop.org/drm/amd/-/issues/5387: here the swapout was > driven by runtime memory pressure — a global OOM followed by > `ttm_global_swapout()` reached from > `amdgpu_gem_fault()`/`ttm_tt_populate()`. > If the mechanism is the same, the defect is not suspend-specific; any > memory-pressure eviction pass can leave the stale endpoint. > > The upstream fix history is worth noting: > > - 3db7d7d < > https://github.com/torvalds/linux/commit/3db7d7d583419f7b1f2e141e36418802dbb25cf8 > > > was the intended fix, but it was applied to the wrong `if`: it changed > `if (ret)` to `if (ret > 0)` after `ttm_resource_try_charge()` in > `ttm_bo_alloc_at_place()` (a separate dmem-charge bypass) and left the > `if (!ret)` after `ttm_tt_swapout()` untouched. A kernel carrying only > `3db7d7d` therefore still has the original bug. > - fcfe647 < > https://github.com/torvalds/linux/commit/fcfe64715b425262af1b36f498f9197f3537ceed > > > applies the intended `if (ret > 0)` at the swapout site and restores the > charge check (`Cc: stable # v7.1+`). > > I checked the source release used for `linux-cachyos 7.2.4-1`: > `cachyos-7.2.4-1` < > https://github.com/CachyOS/linux/releases/tag/cachyos-7.2.4-1>. > It still contains: > > ret = ttm_tt_swapout(bdev, tt, swapout_walk->gfp_flags); > if (!ret) { > ttm_resource_del_bulk_move_unevictable(bo->resource, bo); > ttm_resource_move_to_lru_tail(bo->resource); > } > > The installed `ttm.ko` has the matching `test eax,eax` / `jne` branch, so > the > installed binary also has the old zero-only condition. The CachyOS 7.2.4 > PKGBUILD does not list a TTM bulk-move patch. Version check > (upstream sources): > > - `v6.18.52` does not contain `b2ed01e7ad3d` — it keeps the pre-change > structure: the resource is removed from the bulk_move before the > swapout and re-added conditionally, so it is not affected. > - `v7.2.6` and `v7.2.7` still have `if (!ret)` at this site. The host now > runs `7.2.6-1-cachyos` (booted) and its installed `ttm.ko` still shows > the > same code generation. > - `v7.3-rc4` has the corrected `if (ret > 0)` form. > > The kernel was tainted with `CPU_OUT_OF_SPEC`, `OOT_MODULE`, and > `UNSIGNED_MODULE`; the Oops also set `WARN` and `DIE`. For completeness: > the only out-of-tree module was `v4l2loopback` (a DKMS V4L2 loopback > device) — not in the DRM/TTM path. > > Caveats: I have not boot-tested a fixed kernel on this hardware — no > native kernel A/B was performed. The evidence above is source-level plus > disassembly of the installed `ttm.ko`; `drm/amd#5387` has the live > reproduction data on a different GPU. > > Since `fcfe647` already carries `Cc: stable # v7.1+`, I assume the 7.2.y > backport is queued — `7.2.7` (latest 7.2.y) still carries the bug. Two > things this report adds: (a) an independent hit through a non-hibernation > trigger, on a distribution kernel (CachyOS 7.2.4/7.2.6/7.2.7 all show the > old condition); (b) a heads-up that trees which picked up `3db7d7d` > alone are still buggy — its one-line change landed on the wrong `if`, so > the intended swapout-site fix only exists in `fcfe647` (this > backport-unit warning is now also on public record via a comment on the > CachyOS issue). Does the crash signature above match the dangling-cursor > mechanism? >
