Hi Vadim,

Thanks for confirming — the pos->first/pos->last detail matches what
the journal showed, and the bdev->lru_lock explanation lines up
exactly with the KFD / amdgpu_drm_release hangs I saw afterwards.

The oops details are now on drm/amd#5387 as requested. I'll report
back once 7.2.8 reaches me via the CachyOS package.

Best regards,
Yuu Akaoka (nekomario28)

2026年9月25日(金) 18:16 Vadim Nikitushkin <[email protected]>:

> Hi,
>
> Yes, the signature matches. ttm_resource.c:235 in 7.2.y is the
> WARN_ON(pos->first->bo->base.resv != res->bo->base.resv) in
> ttm_lru_bulk_move_add(): pos->first points at a freed resource, the resv
> comparison reads garbage, and ttm_lru_bulk_move_pos_tail() then
> dereferences the stale pos->last — hence the NULL deref at +0xc9. The oops
> happens under bdev->lru_lock, which is why the KFD worker and
> amdgpu_drm_release then hang on spinlocks.
>
> Hibernation was just my trigger; ttm_device_prepare_hibernation() is a
> global swapout, same ttm_bo_swapout_cb() path as ttm_global_swapout() under
> memory pressure. So yes, any swapout pass can leave the dangling endpoint —
> your report is the first hit I know of on a dGPU and without suspend. If
> you can, please add the oops to drm/amd#5387
> <https://github.com/drm/amd/issues/5387>, it's useful data.
>
> On 7.2.y: the backport is in the 7.2.8-rc1 review series as "[PATCH 7.2
> 436/438] drm/ttm: fix swapped-out resources never leaving their bulk_move
> range" (review closes 25 Sep 14:05 UTC, so 7.2.8 should follow shortly).
> 7.2.y has no ttm_resource_try_charge(), so 3db7d7d can't be cherry-picked
> there at all (Greg's bot reported FAILED); the stable patch is the two
> commits squashed into the single swapout-site hunk. Your point (b) still
> applies to anyone hand-picking from mainline — a tree carrying only 3db7d7d
> is unfixed.
>
> Thanks for the detailed analysis.
>
> Vadim
>
>
>
>
> пт, 25 сент. 2026 г. в 12:03, 赤岡悠 <[email protected]>:
>
>> Hi,
>>
>> I hit a kernel Oops in the AMDGPU/TTM swapout path on a Radeon RX 7800
>> XT. The
>> failure happened immediately after a global OOM condition and was
>> followed by an
>> RCU stall and a host reset.
>>
>> The kernel was `7.2.4-1-cachyos` on CachyOS. The GPU is a Sapphire Navi32
>> / RX
>> 7800 XT (`1002:747e`, subsystem `1da2:475d`). The relevant userspace
>> versions
>> were Mesa `26.2.2-2`, KWin/Plasma `6.7.5-1.1`, and libdrm `2.4.134-1.1`.
>>
>> The first failure in the affected boot was:
>>
>> kswapd0 invoked oom-killer: gfp_mask=0xcc0(GFP_KERNEL), order=0
>> oom-kill: ... global_oom ... task_memcg=<redacted>, task=llama-server
>> Out of memory: Killed process <pid> (llama-server)
>>
>> amdgpu 0000:03:00.0: [drm] *ERROR* Not enough memory for command
>> submission!
>> WARNING: drivers/gpu/drm/ttm/ttm_resource.c:235 at
>> ttm_resource_add_bulk_move+0x124/0x150 [ttm], CPU#15: QSGRenderThread
>> /2457
>> BUG: kernel NULL pointer dereference, address: 0000000000000008
>> Oops: Oops: 0000 [#1] SMP NOPTI
>> RIP: ttm_resource_add_bulk_move+0xc9/0x150 [ttm]
>>
>> The call chain was:
>>
>> ttm_resource_add_bulk_move
>> ttm_resource_alloc
>> ttm_bo_swapout_cb
>> ttm_lru_walk_for_evict
>> ttm_bo_swapout
>> ttm_global_swapout
>> ttm_tt_populate
>> ttm_bo_populate
>> ttm_bo_vm_fault_reserved
>> amdgpu_gem_fault
>>
>> After the Oops, an `amdgpu` KFD cleanup worker and an `amdgpu_drm_release`
>> path were both blocked on TTM-related spinlocks. The journal does not
>> contain
>> `Kernel panic - not syncing`; I am calling this an Oops followed by a
>> lockup or
>> reset, not a confirmed formal panic.
>>
>> This is consistent with the stale `bulk_move` endpoint bug introduced by
>> `b2ed01e7ad3d` ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on
>> swapout
>> failure", v7.1+): `ttm_tt_swapout()` returns the number of pages swapped
>> on
>> success (never zero for a populated ttm), so the
>> `ttm_resource_del_bulk_move_unevictable()` /
>> `ttm_resource_move_to_lru_tail()` pair under `if (!ret)` is skipped on
>> every
>> successful swapout. A swapped-out resource stays inside its BO's bulk_move
>> range while unevictable; a later `ttm_resource_free()` /
>> `ttm_bo_set_bulk_move()` skips it via the `!ttm_resource_unevictable()`
>> guard, leaving a dangling range endpoint. Upstream names
>> `ttm_resource_add_bulk_move()` as one of the places that trips over that
>> cursor (a use-after-free on the dangling endpoint) — exactly the function
>> my NULL dereference hit at `+0xc9`.
>>
>> The trigger on my system differs from the hibernation cycles reported in
>> https://gitlab.freedesktop.org/drm/amd/-/issues/5387: here the swapout
>> was
>> driven by runtime memory pressure — a global OOM followed by
>> `ttm_global_swapout()` reached from
>> `amdgpu_gem_fault()`/`ttm_tt_populate()`.
>> If the mechanism is the same, the defect is not suspend-specific; any
>> memory-pressure eviction pass can leave the stale endpoint.
>>
>> The upstream fix history is worth noting:
>>
>> - 3db7d7d <
>> https://github.com/torvalds/linux/commit/3db7d7d583419f7b1f2e141e36418802dbb25cf8
>> >
>>   was the intended fix, but it was applied to the wrong `if`: it changed
>>   `if (ret)` to `if (ret > 0)` after `ttm_resource_try_charge()` in
>>   `ttm_bo_alloc_at_place()` (a separate dmem-charge bypass) and left the
>>   `if (!ret)` after `ttm_tt_swapout()` untouched. A kernel carrying only
>>   `3db7d7d` therefore still has the original bug.
>> - fcfe647 <
>> https://github.com/torvalds/linux/commit/fcfe64715b425262af1b36f498f9197f3537ceed
>> >
>>   applies the intended `if (ret > 0)` at the swapout site and restores the
>>   charge check (`Cc: stable # v7.1+`).
>>
>> I checked the source release used for `linux-cachyos 7.2.4-1`:
>> `cachyos-7.2.4-1` <
>> https://github.com/CachyOS/linux/releases/tag/cachyos-7.2.4-1>.
>> It still contains:
>>
>> ret = ttm_tt_swapout(bdev, tt, swapout_walk->gfp_flags);
>> if (!ret) {
>>         ttm_resource_del_bulk_move_unevictable(bo->resource, bo);
>>         ttm_resource_move_to_lru_tail(bo->resource);
>> }
>>
>> The installed `ttm.ko` has the matching `test eax,eax` / `jne` branch, so
>> the
>> installed binary also has the old zero-only condition. The CachyOS 7.2.4
>> PKGBUILD does not list a TTM bulk-move patch. Version check
>> (upstream sources):
>>
>> - `v6.18.52` does not contain `b2ed01e7ad3d` — it keeps the pre-change
>>   structure: the resource is removed from the bulk_move before the
>>   swapout and re-added conditionally, so it is not affected.
>> - `v7.2.6` and `v7.2.7` still have `if (!ret)` at this site. The host now
>>   runs `7.2.6-1-cachyos` (booted) and its installed `ttm.ko` still shows
>> the
>>   same code generation.
>> - `v7.3-rc4` has the corrected `if (ret > 0)` form.
>>
>> The kernel was tainted with `CPU_OUT_OF_SPEC`, `OOT_MODULE`, and
>> `UNSIGNED_MODULE`; the Oops also set `WARN` and `DIE`. For completeness:
>> the only out-of-tree module was `v4l2loopback` (a DKMS V4L2 loopback
>> device) — not in the DRM/TTM path.
>>
>> Caveats: I have not boot-tested a fixed kernel on this hardware — no
>> native kernel A/B was performed. The evidence above is source-level plus
>> disassembly of the installed `ttm.ko`; `drm/amd#5387` has the live
>> reproduction data on a different GPU.
>>
>> Since `fcfe647` already carries `Cc: stable # v7.1+`, I assume the 7.2.y
>> backport is queued — `7.2.7` (latest 7.2.y) still carries the bug. Two
>> things this report adds: (a) an independent hit through a non-hibernation
>> trigger, on a distribution kernel (CachyOS 7.2.4/7.2.6/7.2.7 all show the
>> old condition); (b) a heads-up that trees which picked up `3db7d7d`
>> alone are still buggy — its one-line change landed on the wrong `if`, so
>> the intended swapout-site fix only exists in `fcfe647` (this
>> backport-unit warning is now also on public record via a comment on the
>> CachyOS issue). Does the crash signature above match the dangling-cursor
>> mechanism?
>>
>

Reply via email to