Public bug reported:

When the MT7921 firmware stops responding ("driver own failed"), 
mt7921_mac_reset_work
tries to reset the chip. The reset fails ("chip reset failed"), but the reset 
path
keeps using the DMA rings anyway, and the kernel oopses in the mt76 DMA code. 
The
oops happens in a kworker that holds networking locks, so with the default
panic_on_oops=0 the whole machine hangs: the display freezes, and keyboard and 
mouse
stop working. A hard power-off is the only way out.

I captured 6 occurrences in 13 days (2026-09-14 … 2026-09-26) on two kernel 
builds
(7.0.0-31 and 7.0.0-34). efi_pstore saved full traces for 4 of them. All 4 run 
through
the same work item (`Workqueue: mt76 mt7921_mac_reset_work [mt7921_common]`), 
and
"chip reset failed" is logged just before each oops.

Nothing specific triggers the chip hang. It has happened under heavy CPU load 
and on an
idle system 10 minutes after boot. Wi-Fi power saving was enabled 
(NetworkManager
`wifi.powersave = 3`) and PCIe ASPM was on for the card.

There are two issues here:

1. The firmware/chip hangs ("driver own failed" / "Timeout for driver own", 
once per
   second for 10–25 s). That alone would only be a Wi-Fi outage.
2. **The main bug:** the driver does not bail out when the chip reset fails. It 
goes on
   to re-add the vif and send an MCU command through the TX ring, or to clean 
up the TX
   queues, and that causes memory corruption and a kernel oops. A failed reset 
should
   leave the device unusable, not bring down the kernel.

### Variant A — page fault in mt76_dma_add_buf (3 occurrences:
2026-09-22, -25, -26)

Always a write to an unmapped address ending in `...ff0`: a ring descriptor 
slot in a
region that is no longer mapped.

```
mt7921e 0000:03:00.0: driver own failed
mt7921e 0000:03:00.0: Timeout for driver own
mt7921e 0000:03:00.0: chip reset failed
BUG: unable to handle page fault for address: ffffcdd8c0a88ff0
#PF: supervisor write access in kernel mode
#PF: error_code(0x0002) - not-present page
PGD 100000067 P4D 100000067 PUD 10087a067 PMD 102f2c067 PTE 0
Oops: Oops: 0002 [#1] SMP NOPTI
CPU: 14 UID: 0 PID: 90424 Comm: kworker/u64:4 Tainted: P S      W  OE       
7.0.0-34-generic #34-Ubuntu PREEMPT(lazy)
Hardware name: LENOVO 83EG/LNVNB161216, BIOS PJCN04WW 10/28/2023
Workqueue: mt76 mt7921_mac_reset_work [mt7921_common]
RIP: 0010:mt76_dma_add_buf.isra.0+0x90/0x200 [mt76]
Code: 0f b7 4f 2a 89 45 c8 45 31 c9 4c 89 55 b0 eb 3b 89 c1 80 cd 40 44 3b 4d 
c8 0f 44 c1 8b 5d d0 0f b7 ca 41 83 c1 02 48 83 c6 20 <41> 89 18 45 89 78 08
RSP: 0018:ffffcdd8e4cd3b50 EFLAGS: 00010286
RAX: 0000000040400000 RBX: 00000000ffd34290 RCX: 0000000000000000
RDX: 0000000000000000 RSI: ffffcdd8e4cd3bd8 RDI: ffff8a8780024da8
RBP: ffffcdd8e4cd3ba0 R08: ffffcdd8c0a88ff0 R09: 0000000000000002
R10: 00000000004ffe02 R11: 0000000000400000 R12: 0000000000000000
R13: 0000000000000000 R14: 000000000000ffff R15: 0000000000000000
CR2: ffffcdd8c0a88ff0 CR3: 0000000a8ea44000 CR4: 0000000000f50ef0
Call Trace:
 <TASK>
 mt76_dma_tx_queue_skb_raw+0x118/0x1c0 [mt76]
 mt7921_mcu_send_message+0x5a/0x70 [mt7921e]
 mt76_mcu_skb_send_and_get_msg+0x104/0x2b0 [mt76]
 mt76_mcu_send_and_get_msg+0x82/0xc0 [mt76]
 mt76_connac_mcu_uni_add_dev+0x13a/0x1f0 [mt76_connac_lib]
 mt7921_vif_connect_iter+0x47/0xf0 [mt7921_common]
 __iterate_interfaces+0x95/0x130 [mac80211]
 ieee80211_iterate_interfaces+0x3d/0x60 [mac80211]
 mt7921_mac_reset_work+0x13f/0x1f0 [mt7921_common]
 process_one_work+0x1ac/0x3d0
 worker_thread+0x1b8/0x360
 kthread+0xf7/0x130
 ret_from_fork+0x195/0x2a0
 ret_from_fork_asm+0x1a/0x30
 </TASK>
```

The trace from 2026-09-22 (kernel 7.0.0-31, **not** tainted with P — the 
proprietary
NVIDIA module was not loaded in that boot; taint `G S W OE`) is identical down 
to the
offsets, faulting address `ffffcf5b80935ff0`.

### Variant B — WARN in __iommu_dma_unmap, then GPF in
__mt76_tx_complete_skb (2026-09-14, 7.0.0-31)

Same work item, but it fails earlier, in `mt7921e_mac_reset` → 
`mt792x_wpdma_reset` →
`mt76_dma_tx_cleanup`: unmapping buffers that are not mapped, then a 
use-after-free
while completing TX skbs.

```
mt7921e 0000:03:00.0: chip reset failed
WARNING: drivers/iommu/dma-iommu.c:828 at __iommu_dma_unmap+0x15b/0x170, 
CPU#11: kworker/u64:0/51449
Workqueue: mt76 mt7921_mac_reset_work [mt7921_common]
RIP: 0010:__iommu_dma_unmap+0x15b/0x170
Call Trace:
 iommu_dma_unmap_phys+0x5c/0x100
 dma_unmap_phys+0x24a/0x340
 dma_unmap_page_attrs+0x17/0x40
 mt76_dma_tx_cleanup+0x147/0x2c0 [mt76]
 mt792x_wpdma_reset+0x89/0x300 [mt792x_lib]
 mt7921e_mac_reset+0x139/0x3a0 [mt7921e]
 mt7921_mac_reset_work+0x9a/0x1f0 [mt7921_common]
 process_one_work+0x1ac/0x3d0
 ...
---[ end trace ]---
(same WARNING repeated from mt76_dma_tx_cleanup+0x23d)
Oops: general protection fault, probably for non-canonical address 
0x2ba789e787a991e1: 0000 [#1] SMP NOPTI
RIP: 0010:__mt76_tx_complete_skb+0xc0/0x350 [mt76]
Call Trace:
 mt76_connac_tx_complete_skb+0x23/0x50 [mt76_connac_lib]
 mt76_queue_tx_complete+0x28/0x60 [mt76]
 mt76_dma_tx_cleanup+0x1bb/0x2c0 [mt76]
 mt792x_wpdma_reset+0x89/0x300 [mt792x_lib]
 ...
```

Full logs for all three dates are attached (reassembled from efi_pstore; only 
AppArmor
audit lines are removed).

### Steps to reproduce

No deterministic reproducer. Normal use connected to a 5 GHz AP, sometimes 
under heavy
CPU load; the rate is about one crash every 2–3 days. The chip hang starts on 
its own
("driver own failed"), then comes the failed reset, then the oops.

### Expected

If the chip reset fails, the driver marks the device dead and stops touching 
the DMA
rings. Wi-Fi is lost, but the kernel keeps running.

### Actual

Kernel oops in the reset worker. The system freezes completely, because the oops
happens with locks held and panic_on_oops=0.

### Environment

- Ubuntu 26.04.1 LTS
- `Ubuntu 7.0.0-34.34-generic 7.0.14` (also seen on 7.0.0-31-generic)
- Lenovo Legion R7000 APH9 (83EG, board LNVNB161216), BIOS PJCN04WW 10/28/2023, 
AMD Ryzen 7 7840H
- `03:00.0 Network controller [0280]: MEDIATEK Corp. MT7921 802.11ax PCIe 
Wireless Network Adapter [Filogic 330] [14c3:7961]`, Subsystem Lenovo 
`[17aa:e0bc]`
- ASIC revision `79610010`
- Firmware: HW/SW `0x8a108a10` build `20260224110909a`; WM `____010000` build 
`20260224110949`
- linux-firmware `20260319.git217ca6e4.1ubuntu`
- mt7921e srcversion `5DD510B72B6C96B8A78E227` (same in -31 and -34)
- Module params at crash time: `mt7921e.disable_aspm=N`, 
`mt7921_common.disable_clc=N`
- Wi-Fi power save: on (NetworkManager `wifi.powersave = 3`)

### Workaround being tested

- `options mt7921e disable_aspm=Y`
- NetworkManager `wifi.powersave = 2` (off)
- `kernel.panic_on_oops = 1`, `kernel.panic = 10`, so the machine reboots 
instead of
  freezing

Applied on 2026-09-27. I will report whether the chip hang ("driver own 
failed") comes
back. Even if the workaround hides the trigger, the oops after a failed reset 
is still
a driver bug.

Not yet tested on the mainline kernel. I can test a mainline or -proposed build 
if
requested.

ProblemType: Bug
DistroRelease: Ubuntu 26.04
Package: linux-image-7.0.0-34-generic 7.0.0-34.34
ProcVersionSignature: Ubuntu 7.0.0-34.34-generic 7.0.14
Uname: Linux 7.0.0-34-generic x86_64
NonfreeKernelModules: nvidia_modeset nvidia
ApportVersion: 2.34.1-0ubuntu0.1
Architecture: amd64
AudioDevicesInUse:
 USER        PID ACCESS COMMAND
 /dev/snd/controlC1:  qdrin      3098 F.... wireplumber
 /dev/snd/controlC0:  qdrin      3098 F.... wireplumber
 /dev/snd/seq:        qdrin      3094 F.... pipewire
CasperMD5CheckResult: unknown
CurrentDesktop: KDE
Date: Sun Sep 27 10:02:06 2026
InstallationDate: Installed on 2025-10-22 (340 days ago)
InstallationMedia: Kubuntu 25.10 "Questing Quokka" - Release amd64 (20251007)
MachineType: LENOVO 83EG
ProcFB: 0 nvidia-drmdrmfb
ProcKernelCmdLine: BOOT_IMAGE=/boot/vmlinuz-7.0.0-34-generic 
root=/dev/mapper/vg0-lv0 ro quiet splash tsc=reliable
PulseList: Error: command ['pacmd', 'list'] failed with exit code 1: No 
PulseAudio daemon running, or not running as session daemon.
SourcePackage: linux
UpgradeStatus: Upgraded to resolute on 2026-05-17 (133 days ago)
dmi.bios.date: 10/28/2023
dmi.bios.release: 1.4
dmi.bios.vendor: LENOVO
dmi.bios.version: PJCN04WW
dmi.board.asset.tag: NO Asset Tag
dmi.board.name: LNVNB161216
dmi.board.vendor: LENOVO
dmi.board.version: SDK0T76479 WIN
dmi.chassis.asset.tag: NO Asset Tag
dmi.chassis.type: 10
dmi.chassis.vendor: LENOVO
dmi.chassis.version: Legion R7000 APH9
dmi.ec.firmware.release: 1.4
dmi.modalias: 
dmi:bvnLENOVO:bvrPJCN04WW:bd10/28/2023:br1.4:efr1.4:svnLENOVO:pn83EG:pvrLegionR7000APH9:rvnLENOVO:rnLNVNB161216:rvrSDK0T76479WIN:cvnLENOVO:ct10:cvrLegionR7000APH9:skuLENOVO_MT_83EG_BU_idea_FM_LegionR7000APH9:pfaLegionR7000APH9:
dmi.product.family: Legion R7000 APH9
dmi.product.name: 83EG
dmi.product.sku: LENOVO_MT_83EG_BU_idea_FM_Legion R7000 APH9
dmi.product.version: Legion R7000 APH9
dmi.sys.vendor: LENOVO

** Affects: linux (Ubuntu)
     Importance: Undecided
         Status: New


** Tags: amd64 apport-bug resolute wayland-session

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2168654

Title:
  mt7921e: kernel Oops in mt7921_mac_reset_work after "chip reset
  failed" (page fault in mt76_dma_add_buf / GPF in
  __mt76_tx_complete_skb) freezes the system

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2168654/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to