Public bug reported:

[Impact]

Hibernating a machine with an amdgpu GPU leaves TTM's bulk_move cursors
pointing at freed memory, which shows up minutes to hours later -- or at
process exit / reboot after the resume -- as one of:

  * WARN in ttm_lru_bulk_move_add() (dma_resv not held),
  * "list_del corruption" in ttm_resource_move_to_lru_tail(),
  * NULL pointer dereference in ttm_resource_manager_next().

The result is a hung or oopsing GPU driver after resume and, in our case,
a panic on shutdown. On this machine 5 of 18 suspend-then-hibernate cycles
on the stock 7.0.0-31 kernel ended in one of the above.

Hibernation is only the most reliable trigger, because it swaps out every
BO at once. The same path runs under ordinary memory pressure: an
independent report on a Radeon RX 7800 XT (discrete Navi32, kernel 7.2.4)
hit the identical WARN + NULL dereference in ttm_resource_add_bulk_move()
right after a global OOM, with no suspend involved
(https://lore.kernel.org/dri-devel/ca+y9vqnznakcocjwtlea5ts+clm0rpgsa7yre8vopdo4pv-...@mail.gmail.com/).
So any amdgpu system that swaps out GPU memory is exposed, APU or dGPU.

The root cause is in ttm_bo_swapout_cb(). ttm_tt_swapout() returns the
number of pages swapped out on success and a negative error code on
failure; for a populated ttm it never returns zero. Upstream commit
b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout
failure") put the bulk_move bookkeeping under "if (!ret)", so the
ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail()
pair is skipped on every successful swapout. A swapped-out resource then
stays inside its BO's bulk_move range although it is unevictable; when it
is later freed, ttm_resource_del_bulk_move() skips it because of its
!ttm_resource_unevictable() guard and a range endpoint in pos->first /
pos->last is left dangling.

b2ed01e7ad3d is in upstream 7.0.y since 7.0.10, so the Ubuntu 7.0 kernel
carries the regression. Upstream 7.0.y is no longer maintained (it is off
the kernel.org list; last commit 27 June 2026), so the fix will not reach
this kernel through the usual stable updates -- it has to be cherry-picked.

This is drm/amd issue 5387, which has a number of reporters on AMD APUs.

[Fix]

Two mainline commits, both in v7.3-rc4 and both tagged
Cc: [email protected] # v7.1+:

  3db7d7d583419f7b1f2e141e36418802dbb25cf8
      drm/ttm: fix swapped-out resources never leaving their bulk_move range
  fcfe64715b425262af1b36f498f9197f3537ceed
      drm/ttm: apply the swapout bulk_move fix to the intended condition

The first one was applied by the maintainer to the wrong condition (the
"if (ret)" after ttm_resource_try_charge()); the second restored that one
and applied the intended change. ttm_resource_try_charge() does not exist
in 7.0, so neither commit applies on its own and the net effect of the two
is a single hunk in ttm_bo_swapout_cb():

        -               if (!ret) {
        +               if (ret > 0) {

The attached patch is that backport against the Ubuntu 7.0.0-34.34 source
(linux-source-7.0.0; ttm_bo.c is unchanged between -31 and -34), in the
usual "backported from commit ..." form. The identical squashed backport
was accepted by Greg KH for 7.2.y and is released in Linux 7.2.8
(25 September 2026):
  https://lore.kernel.org/stable/[email protected]/
so taking it as-is keeps the Ubuntu kernel in line with upstream stable.

Reviewed upstream by both TTM maintainers (Thomas Hellström, Christian
König).

[Test Plan]

Deterministic check, no crash needed. Count the TTM swapout calls across a
hibernation with ftrace's function profiler:

  # cd /sys/kernel/tracing
  # echo 0 > function_profile_enabled
  # echo 'ttm_tt_swapout ttm_resource_del_bulk_move_unevictable' > 
set_ftrace_filter
  # echo 1 > function_profile_enabled
  <hibernate and resume>
  # grep ttm_ trace_stat/function*

Broken kernel: hundreds of ttm_tt_swapout() calls and zero
ttm_resource_del_bulk_move_unevictable() calls. Fixed kernel: the two
counts track each other.

Crash reproducer: repeat suspend-then-hibernate

  # systemctl suspend-then-hibernate

with a normal desktop session (GNOME/Wayland) loaded, resume, use the
machine, then log out or reboot. On the stock kernel one of the signatures
above appears within roughly 20 cycles. With the fix, the same test ran
clean.

[Where problems could occur]

The change is one line, confined to the swapout path in ttm_bo_swapout_cb()
(drivers/gpu/drm/ttm/ttm_bo.c), and it restores the behaviour that existed
before b2ed01e7ad3d: the resource is taken off its bulk_move range when it
is actually swapped out. It matches the equivalent shrinker check in
1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink() infinite LRU walk on backup
failure"), which tests "lret > 0".

TTM is shared by amdgpu, radeon, nouveau, qxl, vmwgfx and xe, so all of
them are touched, but only the bookkeeping after a successful swapout
changes. If ttm_tt_swapout() ever returned 0 for a populated ttm the
bookkeeping would be skipped exactly as it is today, so there is no new
failure mode; the worst case is slightly different LRU ordering.

[Other Info]

Upstream discussion:
  
https://lore.kernel.org/dri-devel/[email protected]/T/#u
Upstream bug:
  https://gitlab.freedesktop.org/drm/amd/-/issues/5387
7.2.y backport thread (Greg KH's FAILED notice, the backport, his ack):
  https://lore.kernel.org/stable/2026092253-dimness-unethical-2515@gregkh/T/#u
Independent non-hibernation report (RX 7800 XT, OOM-triggered):
  
https://lore.kernel.org/dri-devel/ca+y9vqnznakcocjwtlea5ts+clm0rpgsa7yre8vopdo4pv-...@mail.gmail.com/T/#u

A locally built ttm.ko carrying this change has been in use on the
affected machine since 9 September 2026 (7.0.0-31, then 7.0.0-34) with
no recurrence across 20 hibernation cycles (journal checked for all three
signatures above: zero hits).
--- 
ProblemType: Bug
ApportVersion: 2.34.1-0ubuntu0.1
Architecture: amd64
AudioDevicesInUse:
 USER        PID ACCESS COMMAND
 /dev/snd/controlC1:  cool-t     2256 F.... wireplumber
 /dev/snd/controlC0:  cool-t     2256 F.... wireplumber
 /dev/snd/seq:        cool-t     2238 F.... pipewire
CasperMD5CheckResult: fail
CurrentDesktop: ubuntu:GNOME
DistroRelease: Ubuntu 26.04
HibernationDevice: RESUME=/dev/mapper/ubuntu--vg-ubuntu--lv
InstallationDate: Installed on 2026-08-22 (34 days ago)
InstallationMedia: Ubuntu-Server 26.04 "Resolute Raccoon" - Release amd64 
(20260420.1)
MachineType: ASUS Zenbook 14 UM3406GA
Package: linux (not installed)
ProcFB: 0 amdgpudrmfb
ProcKernelCmdLine: BOOT_IMAGE=/vmlinuz-7.0.0-34-generic 
root=/dev/mapper/ubuntu--vg-ubuntu--lv ro 
resume=/dev/mapper/ubuntu--vg-ubuntu--lv resume_offset=30228480
ProcVersionSignature: Ubuntu 7.0.0-34.34-generic 7.0.14
Tags: resolute wayland-session
Uname: Linux 7.0.0-34-generic x86_64
UpgradeStatus: No upgrade log present (probably fresh install)
UserGroups: adm cdrom dip docker i2c input kvm libvirt lxd plugdev sudo uinput 
users
_MarkForUpload: True
dmi.bios.date: 12/09/2025
dmi.bios.release: 5.35
dmi.bios.vendor: American Megatrends International, LLC.
dmi.bios.version: UM3406GA.302
dmi.board.asset.tag: ATN12345678901234567
dmi.board.name: UM3406GA
dmi.board.vendor: ASUSTeK COMPUTER INC.
dmi.board.version: 1.0
dmi.chassis.asset.tag: No Asset Tag
dmi.chassis.type: 10
dmi.chassis.vendor: ASUSTeK COMPUTER INC.
dmi.chassis.version: 1.0
dmi.ec.firmware.release: 3.3
dmi.modalias: 
dmi:bvnAmericanMegatrendsInternational,LLC.:bvrUM3406GA.302:bd12/09/2025:br5.35:efr3.3:svnASUS:pnZenbook14UM3406GA:pvr1.0:rvnASUSTeKCOMPUTERINC.:rnUM3406GA:rvr1.0:cvnASUSTeKCOMPUTERINC.:ct10:cvr1.0:sku:pfaZenbook14:
dmi.product.family: Zenbook 14
dmi.product.name: Zenbook 14 UM3406GA
dmi.product.version: 1.0
dmi.sys.vendor: ASUS

** Affects: linux (Ubuntu)
     Importance: Undecided
         Status: New


** Tags: apport-collected resolute wayland-session

** Tags added: apport-collected resolute wayland-session

** Description changed:

  [Impact]
  
  Hibernating a machine with an amdgpu GPU leaves TTM's bulk_move cursors
  pointing at freed memory, which shows up minutes to hours later -- or at
  process exit / reboot after the resume -- as one of:
  
    * WARN in ttm_lru_bulk_move_add() (dma_resv not held),
    * "list_del corruption" in ttm_resource_move_to_lru_tail(),
    * NULL pointer dereference in ttm_resource_manager_next().
  
  The result is a hung or oopsing GPU driver after resume and, in our case,
  a panic on shutdown. On this machine 5 of 18 suspend-then-hibernate cycles
  on the stock 7.0.0-31 kernel ended in one of the above.
  
  Hibernation is only the most reliable trigger, because it swaps out every
  BO at once. The same path runs under ordinary memory pressure: an
  independent report on a Radeon RX 7800 XT (discrete Navi32, kernel 7.2.4)
  hit the identical WARN + NULL dereference in ttm_resource_add_bulk_move()
  right after a global OOM, with no suspend involved
  
(https://lore.kernel.org/dri-devel/ca+y9vqnznakcocjwtlea5ts+clm0rpgsa7yre8vopdo4pv-...@mail.gmail.com/).
  So any amdgpu system that swaps out GPU memory is exposed, APU or dGPU.
  
  The root cause is in ttm_bo_swapout_cb(). ttm_tt_swapout() returns the
  number of pages swapped out on success and a negative error code on
  failure; for a populated ttm it never returns zero. Upstream commit
  b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout
  failure") put the bulk_move bookkeeping under "if (!ret)", so the
  ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail()
  pair is skipped on every successful swapout. A swapped-out resource then
  stays inside its BO's bulk_move range although it is unevictable; when it
  is later freed, ttm_resource_del_bulk_move() skips it because of its
  !ttm_resource_unevictable() guard and a range endpoint in pos->first /
  pos->last is left dangling.
  
  b2ed01e7ad3d is in upstream 7.0.y since 7.0.10, so the Ubuntu 7.0 kernel
  carries the regression. Upstream 7.0.y is no longer maintained (it is off
  the kernel.org list; last commit 27 June 2026), so the fix will not reach
  this kernel through the usual stable updates -- it has to be cherry-picked.
  
  This is drm/amd issue 5387, which has a number of reporters on AMD APUs.
  
  [Fix]
  
  Two mainline commits, both in v7.3-rc4 and both tagged
  Cc: [email protected] # v7.1+:
  
    3db7d7d583419f7b1f2e141e36418802dbb25cf8
        drm/ttm: fix swapped-out resources never leaving their bulk_move range
    fcfe64715b425262af1b36f498f9197f3537ceed
        drm/ttm: apply the swapout bulk_move fix to the intended condition
  
  The first one was applied by the maintainer to the wrong condition (the
  "if (ret)" after ttm_resource_try_charge()); the second restored that one
  and applied the intended change. ttm_resource_try_charge() does not exist
  in 7.0, so neither commit applies on its own and the net effect of the two
  is a single hunk in ttm_bo_swapout_cb():
  
        -               if (!ret) {
        +               if (ret > 0) {
  
  The attached patch is that backport against the Ubuntu 7.0.0-34.34 source
  (linux-source-7.0.0; ttm_bo.c is unchanged between -31 and -34), in the
  usual "backported from commit ..." form. The identical squashed backport
  was accepted by Greg KH for 7.2.y and is released in Linux 7.2.8
  (25 September 2026):
    https://lore.kernel.org/stable/[email protected]/
  so taking it as-is keeps the Ubuntu kernel in line with upstream stable.
  
  Reviewed upstream by both TTM maintainers (Thomas Hellström, Christian
  König).
  
  [Test Plan]
  
  Deterministic check, no crash needed. Count the TTM swapout calls across a
  hibernation with ftrace's function profiler:
  
    # cd /sys/kernel/tracing
    # echo 0 > function_profile_enabled
    # echo 'ttm_tt_swapout ttm_resource_del_bulk_move_unevictable' > 
set_ftrace_filter
    # echo 1 > function_profile_enabled
    <hibernate and resume>
    # grep ttm_ trace_stat/function*
  
  Broken kernel: hundreds of ttm_tt_swapout() calls and zero
  ttm_resource_del_bulk_move_unevictable() calls. Fixed kernel: the two
  counts track each other.
  
  Crash reproducer: repeat suspend-then-hibernate
  
    # systemctl suspend-then-hibernate
  
  with a normal desktop session (GNOME/Wayland) loaded, resume, use the
  machine, then log out or reboot. On the stock kernel one of the signatures
  above appears within roughly 20 cycles. With the fix, the same test ran
  clean.
  
  [Where problems could occur]
  
  The change is one line, confined to the swapout path in ttm_bo_swapout_cb()
  (drivers/gpu/drm/ttm/ttm_bo.c), and it restores the behaviour that existed
  before b2ed01e7ad3d: the resource is taken off its bulk_move range when it
  is actually swapped out. It matches the equivalent shrinker check in
  1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink() infinite LRU walk on backup
  failure"), which tests "lret > 0".
  
  TTM is shared by amdgpu, radeon, nouveau, qxl, vmwgfx and xe, so all of
  them are touched, but only the bookkeeping after a successful swapout
  changes. If ttm_tt_swapout() ever returned 0 for a populated ttm the
  bookkeeping would be skipped exactly as it is today, so there is no new
  failure mode; the worst case is slightly different LRU ordering.
  
  [Other Info]
  
  Upstream discussion:
    
https://lore.kernel.org/dri-devel/[email protected]/T/#u
  Upstream bug:
    https://gitlab.freedesktop.org/drm/amd/-/issues/5387
  7.2.y backport thread (Greg KH's FAILED notice, the backport, his ack):
    https://lore.kernel.org/stable/2026092253-dimness-unethical-2515@gregkh/T/#u
  Independent non-hibernation report (RX 7800 XT, OOM-triggered):
    
https://lore.kernel.org/dri-devel/ca+y9vqnznakcocjwtlea5ts+clm0rpgsa7yre8vopdo4pv-...@mail.gmail.com/T/#u
  
  A locally built ttm.ko carrying this change has been in use on the
  affected machine since 9 September 2026 (7.0.0-31, then 7.0.0-34) with
  no recurrence across 20 hibernation cycles (journal checked for all three
  signatures above: zero hits).
+ --- 
+ ProblemType: Bug
+ ApportVersion: 2.34.1-0ubuntu0.1
+ Architecture: amd64
+ AudioDevicesInUse:
+  USER        PID ACCESS COMMAND
+  /dev/snd/controlC1:  cool-t     2256 F.... wireplumber
+  /dev/snd/controlC0:  cool-t     2256 F.... wireplumber
+  /dev/snd/seq:        cool-t     2238 F.... pipewire
+ CasperMD5CheckResult: fail
+ CurrentDesktop: ubuntu:GNOME
+ DistroRelease: Ubuntu 26.04
+ HibernationDevice: RESUME=/dev/mapper/ubuntu--vg-ubuntu--lv
+ InstallationDate: Installed on 2026-08-22 (34 days ago)
+ InstallationMedia: Ubuntu-Server 26.04 "Resolute Raccoon" - Release amd64 
(20260420.1)
+ MachineType: ASUS Zenbook 14 UM3406GA
+ Package: linux (not installed)
+ ProcFB: 0 amdgpudrmfb
+ ProcKernelCmdLine: BOOT_IMAGE=/vmlinuz-7.0.0-34-generic 
root=/dev/mapper/ubuntu--vg-ubuntu--lv ro 
resume=/dev/mapper/ubuntu--vg-ubuntu--lv resume_offset=30228480
+ ProcVersionSignature: Ubuntu 7.0.0-34.34-generic 7.0.14
+ Tags: resolute wayland-session
+ Uname: Linux 7.0.0-34-generic x86_64
+ UpgradeStatus: No upgrade log present (probably fresh install)
+ UserGroups: adm cdrom dip docker i2c input kvm libvirt lxd plugdev sudo 
uinput users
+ _MarkForUpload: True
+ dmi.bios.date: 12/09/2025
+ dmi.bios.release: 5.35
+ dmi.bios.vendor: American Megatrends International, LLC.
+ dmi.bios.version: UM3406GA.302
+ dmi.board.asset.tag: ATN12345678901234567
+ dmi.board.name: UM3406GA
+ dmi.board.vendor: ASUSTeK COMPUTER INC.
+ dmi.board.version: 1.0
+ dmi.chassis.asset.tag: No Asset Tag
+ dmi.chassis.type: 10
+ dmi.chassis.vendor: ASUSTeK COMPUTER INC.
+ dmi.chassis.version: 1.0
+ dmi.ec.firmware.release: 3.3
+ dmi.modalias: 
dmi:bvnAmericanMegatrendsInternational,LLC.:bvrUM3406GA.302:bd12/09/2025:br5.35:efr3.3:svnASUS:pnZenbook14UM3406GA:pvr1.0:rvnASUSTeKCOMPUTERINC.:rnUM3406GA:rvr1.0:cvnASUSTeKCOMPUTERINC.:ct10:cvr1.0:sku:pfaZenbook14:
+ dmi.product.family: Zenbook 14
+ dmi.product.name: Zenbook 14 UM3406GA
+ dmi.product.version: 1.0
+ dmi.sys.vendor: ASUS

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2168554

Title:
  drm/ttm: use-after-free of the bulk_move cursor after hibernation or
  memory-pressure swapout (amdgpu), fixed upstream in v7.3-rc4 / 7.2.8

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2168554/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to