Public bug reported: [Impact]
Hibernating a machine with an amdgpu GPU leaves TTM's bulk_move cursors pointing at freed memory, which shows up minutes to hours later -- or at process exit / reboot after the resume -- as one of: * WARN in ttm_lru_bulk_move_add() (dma_resv not held), * "list_del corruption" in ttm_resource_move_to_lru_tail(), * NULL pointer dereference in ttm_resource_manager_next(). The result is a hung or oopsing GPU driver after resume and, in our case, a panic on shutdown. On this machine 5 of 18 suspend-then-hibernate cycles on the stock 7.0.0-31 kernel ended in one of the above. Hibernation is only the most reliable trigger, because it swaps out every BO at once. The same path runs under ordinary memory pressure: an independent report on a Radeon RX 7800 XT (discrete Navi32, kernel 7.2.4) hit the identical WARN + NULL dereference in ttm_resource_add_bulk_move() right after a global OOM, with no suspend involved (https://lore.kernel.org/dri-devel/ca+y9vqnznakcocjwtlea5ts+clm0rpgsa7yre8vopdo4pv-...@mail.gmail.com/). So any amdgpu system that swaps out GPU memory is exposed, APU or dGPU. The root cause is in ttm_bo_swapout_cb(). ttm_tt_swapout() returns the number of pages swapped out on success and a negative error code on failure; for a populated ttm it never returns zero. Upstream commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure") put the bulk_move bookkeeping under "if (!ret)", so the ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail() pair is skipped on every successful swapout. A swapped-out resource then stays inside its BO's bulk_move range although it is unevictable; when it is later freed, ttm_resource_del_bulk_move() skips it because of its !ttm_resource_unevictable() guard and a range endpoint in pos->first / pos->last is left dangling. b2ed01e7ad3d is in upstream 7.0.y since 7.0.10, so the Ubuntu 7.0 kernel carries the regression. Upstream 7.0.y is no longer maintained (it is off the kernel.org list; last commit 27 June 2026), so the fix will not reach this kernel through the usual stable updates -- it has to be cherry-picked. This is drm/amd issue 5387, which has a number of reporters on AMD APUs. [Fix] Two mainline commits, both in v7.3-rc4 and both tagged Cc: [email protected] # v7.1+: 3db7d7d583419f7b1f2e141e36418802dbb25cf8 drm/ttm: fix swapped-out resources never leaving their bulk_move range fcfe64715b425262af1b36f498f9197f3537ceed drm/ttm: apply the swapout bulk_move fix to the intended condition The first one was applied by the maintainer to the wrong condition (the "if (ret)" after ttm_resource_try_charge()); the second restored that one and applied the intended change. ttm_resource_try_charge() does not exist in 7.0, so neither commit applies on its own and the net effect of the two is a single hunk in ttm_bo_swapout_cb(): - if (!ret) { + if (ret > 0) { The attached patch is that backport against the Ubuntu 7.0.0-34.34 source (linux-source-7.0.0; ttm_bo.c is unchanged between -31 and -34), in the usual "backported from commit ..." form. The identical squashed backport was accepted by Greg KH for 7.2.y and is released in Linux 7.2.8 (25 September 2026): https://lore.kernel.org/stable/[email protected]/ so taking it as-is keeps the Ubuntu kernel in line with upstream stable. Reviewed upstream by both TTM maintainers (Thomas Hellström, Christian König). [Test Plan] Deterministic check, no crash needed. Count the TTM swapout calls across a hibernation with ftrace's function profiler: # cd /sys/kernel/tracing # echo 0 > function_profile_enabled # echo 'ttm_tt_swapout ttm_resource_del_bulk_move_unevictable' > set_ftrace_filter # echo 1 > function_profile_enabled <hibernate and resume> # grep ttm_ trace_stat/function* Broken kernel: hundreds of ttm_tt_swapout() calls and zero ttm_resource_del_bulk_move_unevictable() calls. Fixed kernel: the two counts track each other. Crash reproducer: repeat suspend-then-hibernate # systemctl suspend-then-hibernate with a normal desktop session (GNOME/Wayland) loaded, resume, use the machine, then log out or reboot. On the stock kernel one of the signatures above appears within roughly 20 cycles. With the fix, the same test ran clean. [Where problems could occur] The change is one line, confined to the swapout path in ttm_bo_swapout_cb() (drivers/gpu/drm/ttm/ttm_bo.c), and it restores the behaviour that existed before b2ed01e7ad3d: the resource is taken off its bulk_move range when it is actually swapped out. It matches the equivalent shrinker check in 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink() infinite LRU walk on backup failure"), which tests "lret > 0". TTM is shared by amdgpu, radeon, nouveau, qxl, vmwgfx and xe, so all of them are touched, but only the bookkeeping after a successful swapout changes. If ttm_tt_swapout() ever returned 0 for a populated ttm the bookkeeping would be skipped exactly as it is today, so there is no new failure mode; the worst case is slightly different LRU ordering. [Other Info] Upstream discussion: https://lore.kernel.org/dri-devel/[email protected]/T/#u Upstream bug: https://gitlab.freedesktop.org/drm/amd/-/issues/5387 7.2.y backport thread (Greg KH's FAILED notice, the backport, his ack): https://lore.kernel.org/stable/2026092253-dimness-unethical-2515@gregkh/T/#u Independent non-hibernation report (RX 7800 XT, OOM-triggered): https://lore.kernel.org/dri-devel/ca+y9vqnznakcocjwtlea5ts+clm0rpgsa7yre8vopdo4pv-...@mail.gmail.com/T/#u A locally built ttm.ko carrying this change has been in use on the affected machine since 9 September 2026 (7.0.0-31, then 7.0.0-34) with no recurrence across 20 hibernation cycles (journal checked for all three signatures above: zero hits). --- ProblemType: Bug ApportVersion: 2.34.1-0ubuntu0.1 Architecture: amd64 AudioDevicesInUse: USER PID ACCESS COMMAND /dev/snd/controlC1: cool-t 2256 F.... wireplumber /dev/snd/controlC0: cool-t 2256 F.... wireplumber /dev/snd/seq: cool-t 2238 F.... pipewire CasperMD5CheckResult: fail CurrentDesktop: ubuntu:GNOME DistroRelease: Ubuntu 26.04 HibernationDevice: RESUME=/dev/mapper/ubuntu--vg-ubuntu--lv InstallationDate: Installed on 2026-08-22 (34 days ago) InstallationMedia: Ubuntu-Server 26.04 "Resolute Raccoon" - Release amd64 (20260420.1) MachineType: ASUS Zenbook 14 UM3406GA Package: linux (not installed) ProcFB: 0 amdgpudrmfb ProcKernelCmdLine: BOOT_IMAGE=/vmlinuz-7.0.0-34-generic root=/dev/mapper/ubuntu--vg-ubuntu--lv ro resume=/dev/mapper/ubuntu--vg-ubuntu--lv resume_offset=30228480 ProcVersionSignature: Ubuntu 7.0.0-34.34-generic 7.0.14 Tags: resolute wayland-session Uname: Linux 7.0.0-34-generic x86_64 UpgradeStatus: No upgrade log present (probably fresh install) UserGroups: adm cdrom dip docker i2c input kvm libvirt lxd plugdev sudo uinput users _MarkForUpload: True dmi.bios.date: 12/09/2025 dmi.bios.release: 5.35 dmi.bios.vendor: American Megatrends International, LLC. dmi.bios.version: UM3406GA.302 dmi.board.asset.tag: ATN12345678901234567 dmi.board.name: UM3406GA dmi.board.vendor: ASUSTeK COMPUTER INC. dmi.board.version: 1.0 dmi.chassis.asset.tag: No Asset Tag dmi.chassis.type: 10 dmi.chassis.vendor: ASUSTeK COMPUTER INC. dmi.chassis.version: 1.0 dmi.ec.firmware.release: 3.3 dmi.modalias: dmi:bvnAmericanMegatrendsInternational,LLC.:bvrUM3406GA.302:bd12/09/2025:br5.35:efr3.3:svnASUS:pnZenbook14UM3406GA:pvr1.0:rvnASUSTeKCOMPUTERINC.:rnUM3406GA:rvr1.0:cvnASUSTeKCOMPUTERINC.:ct10:cvr1.0:sku:pfaZenbook14: dmi.product.family: Zenbook 14 dmi.product.name: Zenbook 14 UM3406GA dmi.product.version: 1.0 dmi.sys.vendor: ASUS ** Affects: linux (Ubuntu) Importance: Undecided Status: New ** Tags: apport-collected resolute wayland-session ** Tags added: apport-collected resolute wayland-session ** Description changed: [Impact] Hibernating a machine with an amdgpu GPU leaves TTM's bulk_move cursors pointing at freed memory, which shows up minutes to hours later -- or at process exit / reboot after the resume -- as one of: * WARN in ttm_lru_bulk_move_add() (dma_resv not held), * "list_del corruption" in ttm_resource_move_to_lru_tail(), * NULL pointer dereference in ttm_resource_manager_next(). The result is a hung or oopsing GPU driver after resume and, in our case, a panic on shutdown. On this machine 5 of 18 suspend-then-hibernate cycles on the stock 7.0.0-31 kernel ended in one of the above. Hibernation is only the most reliable trigger, because it swaps out every BO at once. The same path runs under ordinary memory pressure: an independent report on a Radeon RX 7800 XT (discrete Navi32, kernel 7.2.4) hit the identical WARN + NULL dereference in ttm_resource_add_bulk_move() right after a global OOM, with no suspend involved (https://lore.kernel.org/dri-devel/ca+y9vqnznakcocjwtlea5ts+clm0rpgsa7yre8vopdo4pv-...@mail.gmail.com/). So any amdgpu system that swaps out GPU memory is exposed, APU or dGPU. The root cause is in ttm_bo_swapout_cb(). ttm_tt_swapout() returns the number of pages swapped out on success and a negative error code on failure; for a populated ttm it never returns zero. Upstream commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure") put the bulk_move bookkeeping under "if (!ret)", so the ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail() pair is skipped on every successful swapout. A swapped-out resource then stays inside its BO's bulk_move range although it is unevictable; when it is later freed, ttm_resource_del_bulk_move() skips it because of its !ttm_resource_unevictable() guard and a range endpoint in pos->first / pos->last is left dangling. b2ed01e7ad3d is in upstream 7.0.y since 7.0.10, so the Ubuntu 7.0 kernel carries the regression. Upstream 7.0.y is no longer maintained (it is off the kernel.org list; last commit 27 June 2026), so the fix will not reach this kernel through the usual stable updates -- it has to be cherry-picked. This is drm/amd issue 5387, which has a number of reporters on AMD APUs. [Fix] Two mainline commits, both in v7.3-rc4 and both tagged Cc: [email protected] # v7.1+: 3db7d7d583419f7b1f2e141e36418802dbb25cf8 drm/ttm: fix swapped-out resources never leaving their bulk_move range fcfe64715b425262af1b36f498f9197f3537ceed drm/ttm: apply the swapout bulk_move fix to the intended condition The first one was applied by the maintainer to the wrong condition (the "if (ret)" after ttm_resource_try_charge()); the second restored that one and applied the intended change. ttm_resource_try_charge() does not exist in 7.0, so neither commit applies on its own and the net effect of the two is a single hunk in ttm_bo_swapout_cb(): - if (!ret) { + if (ret > 0) { The attached patch is that backport against the Ubuntu 7.0.0-34.34 source (linux-source-7.0.0; ttm_bo.c is unchanged between -31 and -34), in the usual "backported from commit ..." form. The identical squashed backport was accepted by Greg KH for 7.2.y and is released in Linux 7.2.8 (25 September 2026): https://lore.kernel.org/stable/[email protected]/ so taking it as-is keeps the Ubuntu kernel in line with upstream stable. Reviewed upstream by both TTM maintainers (Thomas Hellström, Christian König). [Test Plan] Deterministic check, no crash needed. Count the TTM swapout calls across a hibernation with ftrace's function profiler: # cd /sys/kernel/tracing # echo 0 > function_profile_enabled # echo 'ttm_tt_swapout ttm_resource_del_bulk_move_unevictable' > set_ftrace_filter # echo 1 > function_profile_enabled <hibernate and resume> # grep ttm_ trace_stat/function* Broken kernel: hundreds of ttm_tt_swapout() calls and zero ttm_resource_del_bulk_move_unevictable() calls. Fixed kernel: the two counts track each other. Crash reproducer: repeat suspend-then-hibernate # systemctl suspend-then-hibernate with a normal desktop session (GNOME/Wayland) loaded, resume, use the machine, then log out or reboot. On the stock kernel one of the signatures above appears within roughly 20 cycles. With the fix, the same test ran clean. [Where problems could occur] The change is one line, confined to the swapout path in ttm_bo_swapout_cb() (drivers/gpu/drm/ttm/ttm_bo.c), and it restores the behaviour that existed before b2ed01e7ad3d: the resource is taken off its bulk_move range when it is actually swapped out. It matches the equivalent shrinker check in 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink() infinite LRU walk on backup failure"), which tests "lret > 0". TTM is shared by amdgpu, radeon, nouveau, qxl, vmwgfx and xe, so all of them are touched, but only the bookkeeping after a successful swapout changes. If ttm_tt_swapout() ever returned 0 for a populated ttm the bookkeeping would be skipped exactly as it is today, so there is no new failure mode; the worst case is slightly different LRU ordering. [Other Info] Upstream discussion: https://lore.kernel.org/dri-devel/[email protected]/T/#u Upstream bug: https://gitlab.freedesktop.org/drm/amd/-/issues/5387 7.2.y backport thread (Greg KH's FAILED notice, the backport, his ack): https://lore.kernel.org/stable/2026092253-dimness-unethical-2515@gregkh/T/#u Independent non-hibernation report (RX 7800 XT, OOM-triggered): https://lore.kernel.org/dri-devel/ca+y9vqnznakcocjwtlea5ts+clm0rpgsa7yre8vopdo4pv-...@mail.gmail.com/T/#u A locally built ttm.ko carrying this change has been in use on the affected machine since 9 September 2026 (7.0.0-31, then 7.0.0-34) with no recurrence across 20 hibernation cycles (journal checked for all three signatures above: zero hits). + --- + ProblemType: Bug + ApportVersion: 2.34.1-0ubuntu0.1 + Architecture: amd64 + AudioDevicesInUse: + USER PID ACCESS COMMAND + /dev/snd/controlC1: cool-t 2256 F.... wireplumber + /dev/snd/controlC0: cool-t 2256 F.... wireplumber + /dev/snd/seq: cool-t 2238 F.... pipewire + CasperMD5CheckResult: fail + CurrentDesktop: ubuntu:GNOME + DistroRelease: Ubuntu 26.04 + HibernationDevice: RESUME=/dev/mapper/ubuntu--vg-ubuntu--lv + InstallationDate: Installed on 2026-08-22 (34 days ago) + InstallationMedia: Ubuntu-Server 26.04 "Resolute Raccoon" - Release amd64 (20260420.1) + MachineType: ASUS Zenbook 14 UM3406GA + Package: linux (not installed) + ProcFB: 0 amdgpudrmfb + ProcKernelCmdLine: BOOT_IMAGE=/vmlinuz-7.0.0-34-generic root=/dev/mapper/ubuntu--vg-ubuntu--lv ro resume=/dev/mapper/ubuntu--vg-ubuntu--lv resume_offset=30228480 + ProcVersionSignature: Ubuntu 7.0.0-34.34-generic 7.0.14 + Tags: resolute wayland-session + Uname: Linux 7.0.0-34-generic x86_64 + UpgradeStatus: No upgrade log present (probably fresh install) + UserGroups: adm cdrom dip docker i2c input kvm libvirt lxd plugdev sudo uinput users + _MarkForUpload: True + dmi.bios.date: 12/09/2025 + dmi.bios.release: 5.35 + dmi.bios.vendor: American Megatrends International, LLC. + dmi.bios.version: UM3406GA.302 + dmi.board.asset.tag: ATN12345678901234567 + dmi.board.name: UM3406GA + dmi.board.vendor: ASUSTeK COMPUTER INC. + dmi.board.version: 1.0 + dmi.chassis.asset.tag: No Asset Tag + dmi.chassis.type: 10 + dmi.chassis.vendor: ASUSTeK COMPUTER INC. + dmi.chassis.version: 1.0 + dmi.ec.firmware.release: 3.3 + dmi.modalias: dmi:bvnAmericanMegatrendsInternational,LLC.:bvrUM3406GA.302:bd12/09/2025:br5.35:efr3.3:svnASUS:pnZenbook14UM3406GA:pvr1.0:rvnASUSTeKCOMPUTERINC.:rnUM3406GA:rvr1.0:cvnASUSTeKCOMPUTERINC.:ct10:cvr1.0:sku:pfaZenbook14: + dmi.product.family: Zenbook 14 + dmi.product.name: Zenbook 14 UM3406GA + dmi.product.version: 1.0 + dmi.sys.vendor: ASUS -- You received this bug notification because you are a member of Ubuntu Bugs, which is subscribed to Ubuntu. https://bugs.launchpad.net/bugs/2168554 Title: drm/ttm: use-after-free of the bulk_move cursor after hibernation or memory-pressure swapout (amdgpu), fixed upstream in v7.3-rc4 / 7.2.8 To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2168554/+subscriptions -- ubuntu-bugs mailing list [email protected] https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs
