Hi Vadim, Thanks for confirming — the pos->first/pos->last detail matches what the journal showed, and the bdev->lru_lock explanation lines up exactly with the KFD / amdgpu_drm_release hangs I saw afterwards.
The oops details are now on drm/amd#5387 as requested. I'll report back once 7.2.8 reaches me via the CachyOS package. Best regards, Yuu Akaoka (nekomario28) 2026年9月25日(金) 18:16 Vadim Nikitushkin <[email protected]>: > Hi, > > Yes, the signature matches. ttm_resource.c:235 in 7.2.y is the > WARN_ON(pos->first->bo->base.resv != res->bo->base.resv) in > ttm_lru_bulk_move_add(): pos->first points at a freed resource, the resv > comparison reads garbage, and ttm_lru_bulk_move_pos_tail() then > dereferences the stale pos->last — hence the NULL deref at +0xc9. The oops > happens under bdev->lru_lock, which is why the KFD worker and > amdgpu_drm_release then hang on spinlocks. > > Hibernation was just my trigger; ttm_device_prepare_hibernation() is a > global swapout, same ttm_bo_swapout_cb() path as ttm_global_swapout() under > memory pressure. So yes, any swapout pass can leave the dangling endpoint — > your report is the first hit I know of on a dGPU and without suspend. If > you can, please add the oops to drm/amd#5387 > <https://github.com/drm/amd/issues/5387>, it's useful data. > > On 7.2.y: the backport is in the 7.2.8-rc1 review series as "[PATCH 7.2 > 436/438] drm/ttm: fix swapped-out resources never leaving their bulk_move > range" (review closes 25 Sep 14:05 UTC, so 7.2.8 should follow shortly). > 7.2.y has no ttm_resource_try_charge(), so 3db7d7d can't be cherry-picked > there at all (Greg's bot reported FAILED); the stable patch is the two > commits squashed into the single swapout-site hunk. Your point (b) still > applies to anyone hand-picking from mainline — a tree carrying only 3db7d7d > is unfixed. > > Thanks for the detailed analysis. > > Vadim > > > > > пт, 25 сент. 2026 г. в 12:03, 赤岡悠 <[email protected]>: > >> Hi, >> >> I hit a kernel Oops in the AMDGPU/TTM swapout path on a Radeon RX 7800 >> XT. The >> failure happened immediately after a global OOM condition and was >> followed by an >> RCU stall and a host reset. >> >> The kernel was `7.2.4-1-cachyos` on CachyOS. The GPU is a Sapphire Navi32 >> / RX >> 7800 XT (`1002:747e`, subsystem `1da2:475d`). The relevant userspace >> versions >> were Mesa `26.2.2-2`, KWin/Plasma `6.7.5-1.1`, and libdrm `2.4.134-1.1`. >> >> The first failure in the affected boot was: >> >> kswapd0 invoked oom-killer: gfp_mask=0xcc0(GFP_KERNEL), order=0 >> oom-kill: ... global_oom ... task_memcg=<redacted>, task=llama-server >> Out of memory: Killed process <pid> (llama-server) >> >> amdgpu 0000:03:00.0: [drm] *ERROR* Not enough memory for command >> submission! >> WARNING: drivers/gpu/drm/ttm/ttm_resource.c:235 at >> ttm_resource_add_bulk_move+0x124/0x150 [ttm], CPU#15: QSGRenderThread >> /2457 >> BUG: kernel NULL pointer dereference, address: 0000000000000008 >> Oops: Oops: 0000 [#1] SMP NOPTI >> RIP: ttm_resource_add_bulk_move+0xc9/0x150 [ttm] >> >> The call chain was: >> >> ttm_resource_add_bulk_move >> ttm_resource_alloc >> ttm_bo_swapout_cb >> ttm_lru_walk_for_evict >> ttm_bo_swapout >> ttm_global_swapout >> ttm_tt_populate >> ttm_bo_populate >> ttm_bo_vm_fault_reserved >> amdgpu_gem_fault >> >> After the Oops, an `amdgpu` KFD cleanup worker and an `amdgpu_drm_release` >> path were both blocked on TTM-related spinlocks. The journal does not >> contain >> `Kernel panic - not syncing`; I am calling this an Oops followed by a >> lockup or >> reset, not a confirmed formal panic. >> >> This is consistent with the stale `bulk_move` endpoint bug introduced by >> `b2ed01e7ad3d` ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on >> swapout >> failure", v7.1+): `ttm_tt_swapout()` returns the number of pages swapped >> on >> success (never zero for a populated ttm), so the >> `ttm_resource_del_bulk_move_unevictable()` / >> `ttm_resource_move_to_lru_tail()` pair under `if (!ret)` is skipped on >> every >> successful swapout. A swapped-out resource stays inside its BO's bulk_move >> range while unevictable; a later `ttm_resource_free()` / >> `ttm_bo_set_bulk_move()` skips it via the `!ttm_resource_unevictable()` >> guard, leaving a dangling range endpoint. Upstream names >> `ttm_resource_add_bulk_move()` as one of the places that trips over that >> cursor (a use-after-free on the dangling endpoint) — exactly the function >> my NULL dereference hit at `+0xc9`. >> >> The trigger on my system differs from the hibernation cycles reported in >> https://gitlab.freedesktop.org/drm/amd/-/issues/5387: here the swapout >> was >> driven by runtime memory pressure — a global OOM followed by >> `ttm_global_swapout()` reached from >> `amdgpu_gem_fault()`/`ttm_tt_populate()`. >> If the mechanism is the same, the defect is not suspend-specific; any >> memory-pressure eviction pass can leave the stale endpoint. >> >> The upstream fix history is worth noting: >> >> - 3db7d7d < >> https://github.com/torvalds/linux/commit/3db7d7d583419f7b1f2e141e36418802dbb25cf8 >> > >> was the intended fix, but it was applied to the wrong `if`: it changed >> `if (ret)` to `if (ret > 0)` after `ttm_resource_try_charge()` in >> `ttm_bo_alloc_at_place()` (a separate dmem-charge bypass) and left the >> `if (!ret)` after `ttm_tt_swapout()` untouched. A kernel carrying only >> `3db7d7d` therefore still has the original bug. >> - fcfe647 < >> https://github.com/torvalds/linux/commit/fcfe64715b425262af1b36f498f9197f3537ceed >> > >> applies the intended `if (ret > 0)` at the swapout site and restores the >> charge check (`Cc: stable # v7.1+`). >> >> I checked the source release used for `linux-cachyos 7.2.4-1`: >> `cachyos-7.2.4-1` < >> https://github.com/CachyOS/linux/releases/tag/cachyos-7.2.4-1>. >> It still contains: >> >> ret = ttm_tt_swapout(bdev, tt, swapout_walk->gfp_flags); >> if (!ret) { >> ttm_resource_del_bulk_move_unevictable(bo->resource, bo); >> ttm_resource_move_to_lru_tail(bo->resource); >> } >> >> The installed `ttm.ko` has the matching `test eax,eax` / `jne` branch, so >> the >> installed binary also has the old zero-only condition. The CachyOS 7.2.4 >> PKGBUILD does not list a TTM bulk-move patch. Version check >> (upstream sources): >> >> - `v6.18.52` does not contain `b2ed01e7ad3d` — it keeps the pre-change >> structure: the resource is removed from the bulk_move before the >> swapout and re-added conditionally, so it is not affected. >> - `v7.2.6` and `v7.2.7` still have `if (!ret)` at this site. The host now >> runs `7.2.6-1-cachyos` (booted) and its installed `ttm.ko` still shows >> the >> same code generation. >> - `v7.3-rc4` has the corrected `if (ret > 0)` form. >> >> The kernel was tainted with `CPU_OUT_OF_SPEC`, `OOT_MODULE`, and >> `UNSIGNED_MODULE`; the Oops also set `WARN` and `DIE`. For completeness: >> the only out-of-tree module was `v4l2loopback` (a DKMS V4L2 loopback >> device) — not in the DRM/TTM path. >> >> Caveats: I have not boot-tested a fixed kernel on this hardware — no >> native kernel A/B was performed. The evidence above is source-level plus >> disassembly of the installed `ttm.ko`; `drm/amd#5387` has the live >> reproduction data on a different GPU. >> >> Since `fcfe647` already carries `Cc: stable # v7.1+`, I assume the 7.2.y >> backport is queued — `7.2.7` (latest 7.2.y) still carries the bug. Two >> things this report adds: (a) an independent hit through a non-hibernation >> trigger, on a distribution kernel (CachyOS 7.2.4/7.2.6/7.2.7 all show the >> old condition); (b) a heads-up that trees which picked up `3db7d7d` >> alone are still buggy — its one-line change landed on the wrong `if`, so >> the intended swapout-site fix only exists in `fcfe647` (this >> backport-unit warning is now also on public record via a comment on the >> CachyOS issue). Does the crash signature above match the dangling-cursor >> mechanism? >> >
