Hi,

Yes, the signature matches. ttm_resource.c:235 in 7.2.y is the
WARN_ON(pos->first->bo->base.resv != res->bo->base.resv) in
ttm_lru_bulk_move_add(): pos->first points at a freed resource, the resv
comparison reads garbage, and ttm_lru_bulk_move_pos_tail() then
dereferences the stale pos->last — hence the NULL deref at +0xc9. The oops
happens under bdev->lru_lock, which is why the KFD worker and
amdgpu_drm_release then hang on spinlocks.

Hibernation was just my trigger; ttm_device_prepare_hibernation() is a
global swapout, same ttm_bo_swapout_cb() path as ttm_global_swapout() under
memory pressure. So yes, any swapout pass can leave the dangling endpoint —
your report is the first hit I know of on a dGPU and without suspend. If
you can, please add the oops to drm/amd#5387
<https://github.com/drm/amd/issues/5387>, it's useful data.

On 7.2.y: the backport is in the 7.2.8-rc1 review series as "[PATCH 7.2
436/438] drm/ttm: fix swapped-out resources never leaving their bulk_move
range" (review closes 25 Sep 14:05 UTC, so 7.2.8 should follow shortly).
7.2.y has no ttm_resource_try_charge(), so 3db7d7d can't be cherry-picked
there at all (Greg's bot reported FAILED); the stable patch is the two
commits squashed into the single swapout-site hunk. Your point (b) still
applies to anyone hand-picking from mainline — a tree carrying only 3db7d7d
is unfixed.

Thanks for the detailed analysis.

Vadim




пт, 25 сент. 2026 г. в 12:03, 赤岡悠 <[email protected]>:

> Hi,
>
> I hit a kernel Oops in the AMDGPU/TTM swapout path on a Radeon RX 7800 XT.
> The
> failure happened immediately after a global OOM condition and was followed
> by an
> RCU stall and a host reset.
>
> The kernel was `7.2.4-1-cachyos` on CachyOS. The GPU is a Sapphire Navi32
> / RX
> 7800 XT (`1002:747e`, subsystem `1da2:475d`). The relevant userspace
> versions
> were Mesa `26.2.2-2`, KWin/Plasma `6.7.5-1.1`, and libdrm `2.4.134-1.1`.
>
> The first failure in the affected boot was:
>
> kswapd0 invoked oom-killer: gfp_mask=0xcc0(GFP_KERNEL), order=0
> oom-kill: ... global_oom ... task_memcg=<redacted>, task=llama-server
> Out of memory: Killed process <pid> (llama-server)
>
> amdgpu 0000:03:00.0: [drm] *ERROR* Not enough memory for command
> submission!
> WARNING: drivers/gpu/drm/ttm/ttm_resource.c:235 at
> ttm_resource_add_bulk_move+0x124/0x150 [ttm], CPU#15: QSGRenderThread
> /2457
> BUG: kernel NULL pointer dereference, address: 0000000000000008
> Oops: Oops: 0000 [#1] SMP NOPTI
> RIP: ttm_resource_add_bulk_move+0xc9/0x150 [ttm]
>
> The call chain was:
>
> ttm_resource_add_bulk_move
> ttm_resource_alloc
> ttm_bo_swapout_cb
> ttm_lru_walk_for_evict
> ttm_bo_swapout
> ttm_global_swapout
> ttm_tt_populate
> ttm_bo_populate
> ttm_bo_vm_fault_reserved
> amdgpu_gem_fault
>
> After the Oops, an `amdgpu` KFD cleanup worker and an `amdgpu_drm_release`
> path were both blocked on TTM-related spinlocks. The journal does not
> contain
> `Kernel panic - not syncing`; I am calling this an Oops followed by a
> lockup or
> reset, not a confirmed formal panic.
>
> This is consistent with the stale `bulk_move` endpoint bug introduced by
> `b2ed01e7ad3d` ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout
> failure", v7.1+): `ttm_tt_swapout()` returns the number of pages swapped on
> success (never zero for a populated ttm), so the
> `ttm_resource_del_bulk_move_unevictable()` /
> `ttm_resource_move_to_lru_tail()` pair under `if (!ret)` is skipped on
> every
> successful swapout. A swapped-out resource stays inside its BO's bulk_move
> range while unevictable; a later `ttm_resource_free()` /
> `ttm_bo_set_bulk_move()` skips it via the `!ttm_resource_unevictable()`
> guard, leaving a dangling range endpoint. Upstream names
> `ttm_resource_add_bulk_move()` as one of the places that trips over that
> cursor (a use-after-free on the dangling endpoint) — exactly the function
> my NULL dereference hit at `+0xc9`.
>
> The trigger on my system differs from the hibernation cycles reported in
> https://gitlab.freedesktop.org/drm/amd/-/issues/5387: here the swapout was
> driven by runtime memory pressure — a global OOM followed by
> `ttm_global_swapout()` reached from
> `amdgpu_gem_fault()`/`ttm_tt_populate()`.
> If the mechanism is the same, the defect is not suspend-specific; any
> memory-pressure eviction pass can leave the stale endpoint.
>
> The upstream fix history is worth noting:
>
> - 3db7d7d <
> https://github.com/torvalds/linux/commit/3db7d7d583419f7b1f2e141e36418802dbb25cf8
> >
>   was the intended fix, but it was applied to the wrong `if`: it changed
>   `if (ret)` to `if (ret > 0)` after `ttm_resource_try_charge()` in
>   `ttm_bo_alloc_at_place()` (a separate dmem-charge bypass) and left the
>   `if (!ret)` after `ttm_tt_swapout()` untouched. A kernel carrying only
>   `3db7d7d` therefore still has the original bug.
> - fcfe647 <
> https://github.com/torvalds/linux/commit/fcfe64715b425262af1b36f498f9197f3537ceed
> >
>   applies the intended `if (ret > 0)` at the swapout site and restores the
>   charge check (`Cc: stable # v7.1+`).
>
> I checked the source release used for `linux-cachyos 7.2.4-1`:
> `cachyos-7.2.4-1` <
> https://github.com/CachyOS/linux/releases/tag/cachyos-7.2.4-1>.
> It still contains:
>
> ret = ttm_tt_swapout(bdev, tt, swapout_walk->gfp_flags);
> if (!ret) {
>         ttm_resource_del_bulk_move_unevictable(bo->resource, bo);
>         ttm_resource_move_to_lru_tail(bo->resource);
> }
>
> The installed `ttm.ko` has the matching `test eax,eax` / `jne` branch, so
> the
> installed binary also has the old zero-only condition. The CachyOS 7.2.4
> PKGBUILD does not list a TTM bulk-move patch. Version check
> (upstream sources):
>
> - `v6.18.52` does not contain `b2ed01e7ad3d` — it keeps the pre-change
>   structure: the resource is removed from the bulk_move before the
>   swapout and re-added conditionally, so it is not affected.
> - `v7.2.6` and `v7.2.7` still have `if (!ret)` at this site. The host now
>   runs `7.2.6-1-cachyos` (booted) and its installed `ttm.ko` still shows
> the
>   same code generation.
> - `v7.3-rc4` has the corrected `if (ret > 0)` form.
>
> The kernel was tainted with `CPU_OUT_OF_SPEC`, `OOT_MODULE`, and
> `UNSIGNED_MODULE`; the Oops also set `WARN` and `DIE`. For completeness:
> the only out-of-tree module was `v4l2loopback` (a DKMS V4L2 loopback
> device) — not in the DRM/TTM path.
>
> Caveats: I have not boot-tested a fixed kernel on this hardware — no
> native kernel A/B was performed. The evidence above is source-level plus
> disassembly of the installed `ttm.ko`; `drm/amd#5387` has the live
> reproduction data on a different GPU.
>
> Since `fcfe647` already carries `Cc: stable # v7.1+`, I assume the 7.2.y
> backport is queued — `7.2.7` (latest 7.2.y) still carries the bug. Two
> things this report adds: (a) an independent hit through a non-hibernation
> trigger, on a distribution kernel (CachyOS 7.2.4/7.2.6/7.2.7 all show the
> old condition); (b) a heads-up that trees which picked up `3db7d7d`
> alone are still buggy — its one-line change landed on the wrong `if`, so
> the intended swapout-site fix only exists in `fcfe647` (this
> backport-unit warning is now also on public record via a comment on the
> CachyOS issue). Does the crash signature above match the dangling-cursor
> mechanism?
>

Reply via email to