Hi,
I hit a kernel Oops in the AMDGPU/TTM swapout path on a Radeon RX 7800 XT.
The
failure happened immediately after a global OOM condition and was followed
by an
RCU stall and a host reset.
The kernel was `7.2.4-1-cachyos` on CachyOS. The GPU is a Sapphire Navi32 /
RX
7800 XT (`1002:747e`, subsystem `1da2:475d`). The relevant userspace
versions
were Mesa `26.2.2-2`, KWin/Plasma `6.7.5-1.1`, and libdrm `2.4.134-1.1`.
The first failure in the affected boot was:
kswapd0 invoked oom-killer: gfp_mask=0xcc0(GFP_KERNEL), order=0
oom-kill: ... global_oom ... task_memcg=<redacted>, task=llama-server
Out of memory: Killed process <pid> (llama-server)
amdgpu 0000:03:00.0: [drm] *ERROR* Not enough memory for command submission!
WARNING: drivers/gpu/drm/ttm/ttm_resource.c:235 at
ttm_resource_add_bulk_move+0x124/0x150 [ttm], CPU#15: QSGRenderThread
/2457
BUG: kernel NULL pointer dereference, address: 0000000000000008
Oops: Oops: 0000 [#1] SMP NOPTI
RIP: ttm_resource_add_bulk_move+0xc9/0x150 [ttm]
The call chain was:
ttm_resource_add_bulk_move
ttm_resource_alloc
ttm_bo_swapout_cb
ttm_lru_walk_for_evict
ttm_bo_swapout
ttm_global_swapout
ttm_tt_populate
ttm_bo_populate
ttm_bo_vm_fault_reserved
amdgpu_gem_fault
After the Oops, an `amdgpu` KFD cleanup worker and an `amdgpu_drm_release`
path were both blocked on TTM-related spinlocks. The journal does not
contain
`Kernel panic - not syncing`; I am calling this an Oops followed by a
lockup or
reset, not a confirmed formal panic.
This is consistent with the stale `bulk_move` endpoint bug introduced by
`b2ed01e7ad3d` ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout
failure", v7.1+): `ttm_tt_swapout()` returns the number of pages swapped on
success (never zero for a populated ttm), so the
`ttm_resource_del_bulk_move_unevictable()` /
`ttm_resource_move_to_lru_tail()` pair under `if (!ret)` is skipped on every
successful swapout. A swapped-out resource stays inside its BO's bulk_move
range while unevictable; a later `ttm_resource_free()` /
`ttm_bo_set_bulk_move()` skips it via the `!ttm_resource_unevictable()`
guard, leaving a dangling range endpoint. Upstream names
`ttm_resource_add_bulk_move()` as one of the places that trips over that
cursor (a use-after-free on the dangling endpoint) — exactly the function
my NULL dereference hit at `+0xc9`.
The trigger on my system differs from the hibernation cycles reported in
https://gitlab.freedesktop.org/drm/amd/-/issues/5387: here the swapout was
driven by runtime memory pressure — a global OOM followed by
`ttm_global_swapout()` reached from
`amdgpu_gem_fault()`/`ttm_tt_populate()`.
If the mechanism is the same, the defect is not suspend-specific; any
memory-pressure eviction pass can leave the stale endpoint.
The upstream fix history is worth noting:
- 3db7d7d <
https://github.com/torvalds/linux/commit/3db7d7d583419f7b1f2e141e36418802dbb25cf8
>
was the intended fix, but it was applied to the wrong `if`: it changed
`if (ret)` to `if (ret > 0)` after `ttm_resource_try_charge()` in
`ttm_bo_alloc_at_place()` (a separate dmem-charge bypass) and left the
`if (!ret)` after `ttm_tt_swapout()` untouched. A kernel carrying only
`3db7d7d` therefore still has the original bug.
- fcfe647 <
https://github.com/torvalds/linux/commit/fcfe64715b425262af1b36f498f9197f3537ceed
>
applies the intended `if (ret > 0)` at the swapout site and restores the
charge check (`Cc: stable # v7.1+`).
I checked the source release used for `linux-cachyos 7.2.4-1`:
`cachyos-7.2.4-1` <
https://github.com/CachyOS/linux/releases/tag/cachyos-7.2.4-1>.
It still contains:
ret = ttm_tt_swapout(bdev, tt, swapout_walk->gfp_flags);
if (!ret) {
ttm_resource_del_bulk_move_unevictable(bo->resource, bo);
ttm_resource_move_to_lru_tail(bo->resource);
}
The installed `ttm.ko` has the matching `test eax,eax` / `jne` branch, so
the
installed binary also has the old zero-only condition. The CachyOS 7.2.4
PKGBUILD does not list a TTM bulk-move patch. Version check
(upstream sources):
- `v6.18.52` does not contain `b2ed01e7ad3d` — it keeps the pre-change
structure: the resource is removed from the bulk_move before the
swapout and re-added conditionally, so it is not affected.
- `v7.2.6` and `v7.2.7` still have `if (!ret)` at this site. The host now
runs `7.2.6-1-cachyos` (booted) and its installed `ttm.ko` still shows the
same code generation.
- `v7.3-rc4` has the corrected `if (ret > 0)` form.
The kernel was tainted with `CPU_OUT_OF_SPEC`, `OOT_MODULE`, and
`UNSIGNED_MODULE`; the Oops also set `WARN` and `DIE`. For completeness:
the only out-of-tree module was `v4l2loopback` (a DKMS V4L2 loopback
device) — not in the DRM/TTM path.
Caveats: I have not boot-tested a fixed kernel on this hardware — no
native kernel A/B was performed. The evidence above is source-level plus
disassembly of the installed `ttm.ko`; `drm/amd#5387` has the live
reproduction data on a different GPU.
Since `fcfe647` already carries `Cc: stable # v7.1+`, I assume the 7.2.y
backport is queued — `7.2.7` (latest 7.2.y) still carries the bug. Two
things this report adds: (a) an independent hit through a non-hibernation
trigger, on a distribution kernel (CachyOS 7.2.4/7.2.6/7.2.7 all show the
old condition); (b) a heads-up that trees which picked up `3db7d7d`
alone are still buggy — its one-line change landed on the wrong `if`, so
the intended swapout-site fix only exists in `fcfe647` (this
backport-unit warning is now also on public record via a comment on the
CachyOS issue). Does the crash signature above match the dangling-cursor
mechanism?