Public bug reported:

# fstrim triggers IRQ stack overflow in device-mapper completion on LVM
RAID1 (7.0.0-34-generic)

Suggested target: Ubuntu, source package `linux`, release 26.04
(Resolute).

## Summary

A server running Ubuntu 26.04.1 LTS repeatedly panics and reboots during
storage trimming. It usually restarts on Monday mornings, coinciding
with the weekly `fstrim.timer`. A manual invocation of `fstrim` on
2026-09-28 reproduced the same panic.

The preserved trace reports an IRQ stack guard-page fault in
`bdev_end_io_acct`, with deeply recursive device-mapper I/O completion
above RAID1 and NVMe completion. Each of four saved crashes contains 67
nested `__dm_io_complete` frames. The configured device tree has only a
few layers, not 67 nested devices.

The precise defect, first affected kernel, and fixed version have not
been identified. This report does not claim a bisected regression or an
exact duplicate of an existing bug.

## Affected system

- Ubuntu 26.04.1 LTS (Resolute Raccoon), x86-64.
- Manually reproduced on `7.0.0-34-generic #34-Ubuntu`, untainted, 
`PREEMPT(lazy)` in the panic.
- Boot log identifies this build as `Ubuntu 7.0.0-34.34-generic 7.0.14`, built 
2026-09-02 with GCC 15.2.0 and GNU ld 2.46.
- An earlier matching panic occurred on `7.0.0-31-generic #31-Ubuntu`, also 
untainted.
- Hardware: Supermicro Super Server / X11SCH-LN4F; BIOS 2.9, date reported by 
DMI as `05/07/2026`.
- CPU: Intel Xeon E-2288G @ 3.70 GHz.
- Secure Boot enabled.
- Relevant kernel parameters: `panic=30 kernel.sysrq=1 crashkernel=4G`.

The `/storage` filesystem is XFS on LVM RAID1. The data path is:

```text
XFS /storage
  /dev/mapper/vg0_storage-lv--0_storage (dm-13, LVM raid1)
    rimage_0 (linear LV) -> /dev/mapper/dm_crypt-3 -> NVMe partition (259:3)
    rimage_1 (linear LV) -> /dev/mapper/dm_crypt-2 -> NVMe partition (259:1)
```

Each RAID member also has a separate linear `rmeta` LV on its
corresponding dm-crypt device. The supplied LVM listing shows no
snapshots. The system also has an XFS root filesystem on a separate LVM
RAID1 volume over encrypted SATA devices and an XFS `/storage-local`
filesystem on another encrypted SATA device. The crash trace
specifically includes NVMe completion.

## Reproduction and observed result

On the existing affected installation, running `fstrim` manually caused
a panic and reboot. The resulting pstore record is dated 2026-09-28
04:10:11 UTC. The exact command-line options and mountpoint for that
manual invocation have not yet been confirmed; they should not be
inferred from the storage layout above.

There has been one reported manual reproduction. No clean-install
reproducer, repeated manual success/failure count, or upstream/mainline
kernel test is available.

The scheduled service uses:

```text
/sbin/fstrim --listed-in /etc/fstab:/proc/self/mountinfo --verbose 
--quiet-unsupported
```

Its timer has `OnCalendar=weekly`, `AccuracySec=1h`,
`RandomizedDelaySec=100min`, and `Persistent=true`. A captured timer
listing shows its last trigger at **2026-09-28 02:24:39 UTC**, 17
seconds before the persistent record timestamp for the second morning
panic. The supplied logs do not independently preserve that service
invocation's start/completion messages.

Expected: trimming completes, or returns an ordinary error, while the
system remains operational.

Actual: IRQ stack exhaustion, a fatal exception in interrupt context,
and a reboot. `panic=30` is configured. The next boot performs XFS log
recovery.

## Panic signature

The following is an explicitly abbreviated excerpt from the manual-test
trace. The attached full trace retains the repeated frames, register
values, module list, and remaining interrupt/task frames.

```text
BUG: IRQ stack guard page was hit at ffffcb1380298ff8
  (stack is ffffcb1380299000..ffffcb138029d000)
Oops: stack guard page: 0000 [#1] SMP NOPTI
CPU: 5 UID: 0 PID: 0 Comm: swapper/5 Not tainted
  7.0.0-34-generic #34-Ubuntu PREEMPT(lazy)
RIP: 0010:bdev_end_io_acct+0x6e/0x280
Call Trace:
 <IRQ>
 dm_io_acct+0x70/0x160
 __dm_io_complete+0x1b3/0x330
 dm_io_dec_pending+0x3c/0xa0
 clone_endio+0x94/0x1d0
 bio_endio+0x15e/0x210
 __dm_io_complete+0x168/0x330
 dm_io_dec_pending+0x3c/0xa0
 clone_endio+0x94/0x1d0
 bio_endio+0x15e/0x210
 [many repeated device-mapper completion frames omitted]
 md_end_clone_io+0x4f/0x140
 bio_endio+0x15e/0x210
 raid_end_bio_io+0x49/0x190 [raid1]
 r1_bio_write_done.part.0+0x3e/0x50 [raid1]
 raid1_end_write_request+0x122/0x3f0 [raid1]
 [additional device-mapper completion frames omitted]
 blk_mq_end_request_batch+0x142/0x680
 nvme_pci_complete_batch+0x57/0x70 [nvme]
 nvme_irq+0x87/0xa0 [nvme]
 [remaining frames omitted]
Kernel panic - not syncing: Fatal exception in interrupt
```

The IRQ stack bounds span 16 KiB. In each record's first panic section,
the interrupt trace contains 67 occurrences each of `__dm_io_complete`,
`dm_io_dec_pending`, and `clone_endio`, plus 69 of `bio_endio`.
Duplicate Oops copies in the pstore output were excluded from these
counts. All three September 28 crashes have the same sequence of
resolved interrupt call-frame names after ignoring offsets and `?`
entries.

This points to excessive recursion in the completion path. It does not
by itself establish the cause of the long completion chain, an NVMe
hardware fault, or a defect specifically in the accounting function
where the stack runs out.

## Saved occurrences

| Persistent record timestamp (UTC) | Kernel | Record ID | Context |
|---|---|---|---|
| 2026-09-21 00:06:16 | 7.0.0-31-generic | 7687773172422148104 | Earlier Monday 
panic |
| 2026-09-28 00:56:06 | 7.0.0-34-generic | 7690383610594983944 | First 
Monday-morning panic |
| 2026-09-28 02:24:56 | 7.0.0-34-generic | 7690406502770671624 | Second 
Monday-morning panic |
| 2026-09-28 04:10:11 | 7.0.0-34-generic | 7690433621194178568 | Manual fstrim 
reproduction |

These times come from archived `dmesg-erst-*` file metadata and may
follow the initial fault by a few seconds. The first fault in the
manual-test trace is at uptime 6242.222617 seconds.

## Available diagnostics and testing limits

Pstore/ERST preserved all four crashes. No vmcore is included in the
available evidence; boot logs show kdump failing with `kexec_file_load
failed: Required key not available` while Secure Boot is enabled.

Disabling `fstrim.timer` and avoiding manual trimming has been
recommended as temporary trigger avoidance. Its effectiveness over a
subsequent Monday has not yet been established. No known-good rollback
kernel is identified; both recorded Ubuntu kernel versions are affected.
No bisect or upstream/mainline reproduction has been performed.

## Attachments

- `zephir-manual-fstrim-crash-trace.txt`: complete first panic section for the 
manual reproduction, omitting duplicate Oops copies.
- `zephir-kernel-bug-evidence.tar.gz`: four original combined pstore 
`dmesg.txt` records, the full manual-test excerpt, storage layout, fstrim 
configuration/timing, selected boot diagnostics, and a provenance/checksum 
manifest.

The original `zephir-pstore-after-test.tar.gz` archive is retained
separately if the raw ERST fragments are needed.

## Related discussion, not a confirmed duplicate

[Recursive bio completion stack-overflow discussion, April
2016](https://lists.openwall.net/linux-kernel/2016/04/28/25) describes
the same broad failure class, but does not establish this fstrim/LVM
RAID1 reproducer or provide a verified fix for the affected Ubuntu
kernels.

ProblemType: Bug
DistroRelease: Ubuntu 26.04
Package: linux-image-7.0.0-34-generic 7.0.0-34.34
ProcVersionSignature: Ubuntu 7.0.0-34.34-generic 7.0.14
Uname: Linux 7.0.0-34-generic x86_64
AlsaVersion: Advanced Linux Sound Architecture Driver Version k7.0.0-34-generic.
AplayDevices: Error: [Errno 2] No such file or directory: 'aplay'
ApportVersion: 2.34.1-0ubuntu0.1
Architecture: amd64
ArecordDevices: Error: [Errno 2] No such file or directory: 'arecord'
AudioDevicesInUse: Error: command ['fuser', '-v', '/dev/snd/by-path', 
'/dev/snd/controlC0', '/dev/snd/hwC0D0', '/dev/snd/pcmC0D9p', 
'/dev/snd/pcmC0D8p', '/dev/snd/pcmC0D7p', '/dev/snd/pcmC0D3p', '/dev/snd/seq', 
'/dev/snd/timer'] failed with exit code 1:
CRDA: N/A
Card0.Amixer.info: Error: [Errno 2] No such file or directory: 'amixer'
Card0.Amixer.values: Error: [Errno 2] No such file or directory: 'amixer'
CasperMD5CheckResult: pass
CloudArchitecture: x86_64
CloudID: none
CloudName: none
CloudPlatform: none
CloudSubPlatform: config
Date: Tue Sep 29 04:52:25 2026
InstallationDate: Installed on 2022-05-01 (1612 days ago)
InstallationMedia: Ubuntu-Server 22.04 LTS "Jammy Jellyfish" - Release amd64 
(20220421)
Lsusb:
 Bus 001 Device 001: ID 1d6b:0002 Linux Foundation 2.0 root hub
 Bus 001 Device 002: ID 0557:7000 ATEN International Co., Ltd Hub
 Bus 001 Device 003: ID 0557:2419 ATEN International Co., Ltd Virtual 
mouse/keyboard device
 Bus 002 Device 001: ID 1d6b:0003 Linux Foundation 3.0 root hub
Lsusb-t:
 /:  Bus 001.Port 001: Dev 001, Class=root_hub, Driver=xhci_hcd/16p, 480M
     |__ Port 014: Dev 002, If 0, Class=Hub, Driver=hub/4p, 480M
         |__ Port 001: Dev 003, If 0, Class=Human Interface Device, 
Driver=usbhid, 1.5M
         |__ Port 001: Dev 003, If 1, Class=Human Interface Device, 
Driver=usbhid, 1.5M
 /:  Bus 002.Port 001: Dev 001, Class=root_hub, Driver=xhci_hcd/10p, 10000M
MachineType: Supermicro Super Server
ProcEnviron:
 LANG=en_US.UTF-8
 PATH=(custom, no user)
 SHELL=/bin/bash
 TERM=xterm-256color
 XDG_RUNTIME_DIR=<set>
ProcFB: 0 astdrmfb
ProcKernelCmdLine: root=/dev/mapper/vg0_root-lv--0_root ro  ipv6.disable=1 
panic=30 kernel.sysrq=1  crashkernel=4G
RfKill:
 
SourcePackage: linux
UpgradeStatus: Upgraded to resolute on 2026-06-06 (115 days ago)
dmi.bios.date: 05/07/2026
dmi.bios.release: 5.13
dmi.bios.vendor: American Megatrends Inc.
dmi.bios.version: 2.9
dmi.board.asset.tag: To be filled by O.E.M.
dmi.board.name: X11SCH-LN4F
dmi.board.vendor: Supermicro
dmi.board.version: 1.01
dmi.chassis.asset.tag: To be filled by O.E.M.
dmi.chassis.type: 17
dmi.chassis.vendor: Supermicro
dmi.chassis.version: 0123456789
dmi.modalias: 
dmi:bvnAmericanMegatrendsInc.:bvr2.9:bd05/07/2026:br5.13:svnSupermicro:pnSuperServer:pvr0123456789:rvnSupermicro:rnX11SCH-LN4F:rvr1.01:cvnSupermicro:ct17:cvr0123456789:skuTobefilledbyO.E.M.:pfaTobefilledbyO.E.M.:
dmi.product.family: To be filled by O.E.M.
dmi.product.name: Super Server
dmi.product.sku: To be filled by O.E.M.
dmi.product.version: 0123456789
dmi.sys.vendor: Supermicro

** Affects: linux (Ubuntu)
     Importance: Undecided
         Status: New


** Tags: amd64 apport-bug resolute

** Attachment added: "zephir-manual-fstrim-crash-trace.txt"
   
https://bugs.launchpad.net/bugs/2169515/+attachment/6005741/+files/zephir-manual-fstrim-crash-trace.txt

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2169515

Title:
  # fstrim triggers IRQ stack overflow in device-mapper completion on
  LVM RAID1 (7.0.0-34-generic)

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2169515/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to