LP#2167595 - field feedback on the v2 patch
(lp2167595_mt7921u_usb_disconnect_deadlock_v2.patch, sha256 c83bc45d...4fefd1)
== Environment ==
Ubuntu 26.04.1 LTS, kernel 7.0.0-34-generic (Ubuntu 7.0.0-34.34, upstream
7.0.14)
Adapter: MediaTek MT7921AU, USB ID 0e8d:7961, driver mt7921u, high-speed (USB
2.0) port
Hardware: Dell Precision 3650 Tower, BIOS 1.48.0
v2 built out-of-tree from linux-source-7.0.0 (7.0.0-34.34), installed into
/lib/modules/7.0.0-34-generic/updates/ and in daily use since 2026-09-26.
Kernel taint is O+E only (out-of-tree unsigned modules: these plus
VirtualBox).
Logs below are sanitized: hostname, BSSID and adapter MAC replaced.
== 1. Real-world validation: the deadlock is gone ==
On 2026-09-26 the adapter failed exactly the way it did in the original report:
8 consecutive "vendor request ... failed:-110" timeouts, Bluetooth on the same
chip timing out as well, then a USB reset loop from which the device never
returned (descriptor read errors -110, "device not accepting address" -62,
finally repeated re-enumeration attempts that all failed).
Timeline (local time, 26.09.2026):
16:29:28 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d02c failed:-110
16:29:31 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d054 failed:-110
16:29:34 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d058 failed:-110
16:29:38 kernel: mt7921u 1-9:1.3: vendor request req:63 off:53b8 failed:-110
16:29:40 kernel: Bluetooth: hci0: Opcode 0x0401 failed: -110
16:29:40 kernel: Bluetooth: hci0: command 0x0401 tx timeout
16:29:41 kernel: mt7921u 1-9:1.3: vendor request req:63 off:53c4 failed:-110
16:29:44 kernel: mt7921u 1-9:1.3: vendor request req:66 off:53c4 failed:-110
16:29:47 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d02c failed:-110
16:29:50 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d054 failed:-110
16:29:51 kernel: Bluetooth: hci0: Failed to write uhw reg(-110)
16:29:53 kernel: wlx00c0cab8f19c: deauthenticating from [BSSID] by local choice
(Reason: 3=DEAUTH_LEAVING)
16:30:07 wpa_supplicant[2255]: wlx00c0cab8f19c: CTRL-EVENT-DISCONNECTED
bssid=[BSSID] reason=3 locally_generated=1
16:30:07 wpa_supplicant[2255]: wlx00c0cab8f19c: Added BSSID [BSSID] into ignore
list, ignoring for 10 seconds
16:30:08 kernel: wlx00c0cab8f19c: failed to remove key (1, ff:ff:ff:ff:ff:ff)
from hardware (-110)
16:30:09 kernel: wlx00c0cab8f19c: failed to remove key (2, ff:ff:ff:ff:ff:ff)
from hardware (-110)
16:30:11 kernel: wlx00c0cab8f19c: failed to remove key (4, ff:ff:ff:ff:ff:ff)
from hardware (-110)
16:30:12 kernel: wlx00c0cab8f19c: failed to remove key (5, ff:ff:ff:ff:ff:ff)
from hardware (-110)
16:30:13 kernel: mt7921u 1-9:1.3: timed out waiting for pending tx
16:30:13 kernel: snd_soc_acpi_intel_sdca_quirks soundwire_generic_allocation
snd_soc_sdw_utils snd_soc_acpi intel_rapl_msr soundwire_bus in
16:30:14 NetworkManager[3140]: device (wlx00c0cab8f19c): state change:
activated -> unmanaged (reason 'unmanaged-link-not-init', managed-typ
16:30:14 NetworkManager[3140]: dhcp4 (wlx00c0cab8f19c): canceled DHCP
transaction
16:30:14 NetworkManager[3140]: dhcp4 (wlx00c0cab8f19c): activation: beginning
transaction (timeout in 45 seconds)
16:30:14 NetworkManager[3140]: dhcp4 (wlx00c0cab8f19c): state changed no lease
16:30:14 ModemManager[2292]: <msg> [base-manager] port wlx00c0cab8f19c released
by device '/sys/devices/pci0000:00/0000:00:14.0/usb1/1-9'
16:30:14 wpa_supplicant[2255]: wlx00c0cab8f19c: PMKSA-CACHE-REMOVED [BSSID] 0
16:30:14 wpa_supplicant[2255]: wlx00c0cab8f19c: CTRL-EVENT-DSCP-POLICY clear_all
16:30:14 wpa_supplicant[2255]: wlx00c0cab8f19c: Removed BSSID [BSSID] from
ignore list (clear)
16:30:14 wpa_supplicant[2255]: wlx00c0cab8f19c: CTRL-EVENT-DSCP-POLICY clear_all
16:30:14 wpa_supplicant[2255]: nl80211: deinit ifname=wlx00c0cab8f19c
disabled_11b_rates=0
16:30:14 kernel: usb 1-9: reset high-speed USB device number 3 using xhci_hcd
16:30:20 kernel: usb 1-9: device descriptor read/64, error -110
16:30:24 systemd[1]: NetworkManager-dispatcher.service: Deactivated
successfully.
16:30:35 kernel: usb 1-9: device descriptor read/64, error -110
16:30:36 kernel: usb 1-9: reset high-speed USB device number 3 using xhci_hcd
16:30:41 kernel: usb 1-9: device descriptor read/64, error -110
16:30:57 kernel: usb 1-9: device descriptor read/64, error -110
16:30:57 kernel: usb 1-9: reset high-speed USB device number 3 using xhci_hcd
16:31:08 kernel: usb 1-9: device not accepting address 3, error -62
16:31:08 kernel: usb 1-9: reset high-speed USB device number 3 using xhci_hcd
16:31:20 kernel: usb 1-9: device not accepting address 3, error -62
16:31:20 kernel: usb 1-9: USB disconnect, device number 3
16:31:20 kernel: usb 1-9: new high-speed USB device number 5 using xhci_hcd
16:31:25 kernel: usb 1-9: device descriptor read/64, error -110
16:31:41 kernel: usb 1-9: device descriptor read/64, error -110
16:31:41 kernel: usb 1-9: new high-speed USB device number 6 using xhci_hcd
16:31:47 kernel: usb 1-9: device descriptor read/64, error -110
16:32:03 kernel: usb 1-9: device descriptor read/64, error -110
16:32:03 kernel: usb 1-9: new high-speed USB device number 7 using xhci_hcd
16:32:14 kernel: usb 1-9: device not accepting address 7, error -62
16:32:14 kernel: usb 1-9: new high-speed USB device number 8 using xhci_hcd
16:32:26 kernel: usb 1-9: device not accepting address 8, error -62
16:45:01 systemd[7730]: Reached target shutdown.target - Shutdown.
16:45:04 systemd[1]: NetworkManager-wait-online.service: Deactivated
successfully.
16:45:05 systemd[1]: NetworkManager.service: Deactivated successfully.
16:45:09 systemd[1]: Reached target shutdown.target - System Shutdown.
16:45:09 systemd[1]: Reached target poweroff.target - System Power Off.
16:45:09 systemd-shutdown[1]: Syncing filesystems and block devices.
16:45:09 systemd-shutdown[1]: Sending SIGTERM to remaining processes...
16:45:09 systemd-journald[1051]: Journal stopped
Outcome with v2:
* No task ever blocked on rtnl_lock. Zero "blocked for more than N seconds"
in the whole boot.
* The system stayed fully responsive for the remaining 15 minutes of the
session (only WiFi was gone) and then shut down normally: the full
shutdown sequence took 8 seconds, from 16:45:01 to "Journal stopped" at
16:45:09.
* No "chip reset failed", no "rx urb mismatch".
For comparison, the same hardware failure on stock/v1 modules (2026-09-15,
the original report) left NetworkManager, ip, and 6 other tasks in D state
for more than 122 seconds, all waiting for rtnl_lock held by a kworker in
mt792xu_disconnect -> mt76_unregister_device -> mt7921_abort_roc, and the
machine had to be powered off with the button.
So the core goal of the patch is confirmed in the field, not just in
tests.
== 2. New finding: the tx_worker disable in mt792xu_disconnect is redundant
and fires a WARNING on the dead-device path ==
Alongside the (harmless) "timed out waiting for pending tx" message, v2
produced a one-shot WARNING:
16:30:13 mt7921u 1-9:1.3: timed out waiting for pending tx
16:30:13 ------------[ cut here ]------------
16:30:13 WARNING: kernel/kthread.c:722 at kthread_park+0x8c/0xc0, CPU#7:
kworker/7:2/100058
16:30:13 CPU: 7 UID: 0 PID: 100058 Comm: kworker/7:2 Tainted: G OE
7.0.0-34-generic #34-Ubuntu PREEMPT(full)
16:30:13 Tainted: [O]=OOT_MODULE, [E]=UNSIGNED_MODULE
16:30:13 Hardware name: Dell Inc. Precision 3650 Tower/0NDYHG, BIOS 1.48.0
05/27/2026
16:30:13 Workqueue: events __usb_queue_reset_device
16:30:13 FS: 0000000000000000(0000) GS:ffff8de34f77f000(0000)
knlGS:0000000000000000
16:30:13 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
16:30:13 Call Trace:
16:30:13 <TASK>
16:30:13 ? mt76u_stop_tx.cold+0x11c/0x180 [mt76_usb]
16:30:13 ? __pfx_autoremove_wake_function+0x10/0x10
16:30:13 mt792xu_stop+0x1a/0x40 [mt792x_usb]
16:30:13 drv_stop+0x50/0x160 [mac80211]
16:30:13 ieee80211_stop_device+0x7f/0x90 [mac80211]
16:30:13 ieee80211_do_stop+0x6a2/0xb20 [mac80211]
16:30:13 ? _raw_spin_lock_irqsave+0xe/0x20
16:30:13 ? packet_notifier+0x85/0x280
16:30:13 ieee80211_stop+0x63/0xf0 [mac80211]
16:30:13 __dev_close_many+0xb2/0x220
16:30:13 netif_close_many+0xb7/0x1b0
16:30:13 netif_close+0x70/0xa0
16:30:13 dev_close+0x38/0xb0
16:30:13 cfg80211_shutdown_all_interfaces+0x50/0x100 [cfg80211]
16:30:13 ieee80211_remove_interfaces+0x47/0x220 [mac80211]
16:30:13 ieee80211_unregister_hw+0x4a/0x140 [mac80211]
16:30:13 mt76_unregister_device+0x63/0x80 [mt76]
16:30:13 mt792xu_disconnect+0xea/0x150 [mt792x_usb]
16:30:13 usb_unbind_interface+0x9b/0x2c0
16:30:13 device_remove+0x68/0x80
16:30:13 device_release_driver_internal+0x1fb/0x260
16:30:13 device_release_driver+0x12/0x20
16:30:13 usb_forced_unbind_intf+0x96/0xe0
16:30:13 ? usb_autoresume_device+0x1e/0x70
16:30:13 usb_reset_device+0xf4/0x300
16:30:13 __usb_queue_reset_device+0x3b/0x60
16:30:13 process_one_work+0x1ac/0x3d0
16:30:13 worker_thread+0x1b8/0x360
16:30:13 ? _raw_spin_lock_irqsave+0xe/0x20
16:30:13 ? __pfx_worker_thread+0x10/0x10
16:30:13 kthread+0xf7/0x130
16:30:13 ? __pfx_kthread+0x10/0x10
16:30:13 ret_from_fork+0x195/0x2a0
16:30:13 ? __pfx_kthread+0x10/0x10
16:30:13 ? __pfx_kthread+0x10/0x10
16:30:13 ret_from_fork_asm+0x1a/0x30
16:30:13 </TASK>
16:30:13 ---[ end trace 0000000000000000 ]---
Analysis:
1. v2 parks tx_worker early in mt792xu_disconnect():
mt76_worker_disable(&dev->mt76.tx_worker);
2. Later in the same function, mt76_unregister_device() ->
ieee80211_unregister_hw() -> ... -> drv_stop() -> mt792xu_stop() ->
mt76u_stop_tx().
3. mt76u_stop_tx() (drivers/net/wireless/mediatek/mt76/usb.c:994) waits
HZ/5 for pending TX to drain. On timeout it logs the message, kills the
TX URBs and calls mt76_worker_disable(&dev->tx_worker) - a SECOND park of
the same worker. kthread_park() then hits
WARN_ON_ONCE(test_bit(KTHREAD_SHOULD_PARK, &kthread->flags))
(kernel/kthread.c:722) and returns -EBUSY.
4. At the end of that same branch mt76u_stop_tx() calls
mt76_worker_enable(&dev->tx_worker), i.e. it UNPARKS the worker that v2
deliberately stopped - so the patch line loses its effect precisely in
the case it was meant to cover.
Suggestion: drop
mt76_worker_disable(&dev->mt76.tx_worker);
from mt792xu_disconnect(). mt76u_stop_tx() already quiesces TX during
unregister and pairs its own park/unpark correctly. The line is not part of
the deadlock fix (flags + worker cancellation + the abort_roc early return
are), and on the dead-device path it is actively counterproductive.
Why lab testing does not catch this: the frame in the trace is
mt76u_stop_tx.cold, i.e. the unlikely branch. With a healthy adapter the TX
queues drain far below the 200 ms timeout, so that branch never executes.
On 2026-09-24 we ran three consecutive rmmod/insmod cycles plus an
unplug-during-transfer test on v2 and saw no WARNING at all; the two MCU
timeout messages were the only output. It took a genuine adapter failure,
with TX still queued at disconnect time, to reach it.
Impact: cosmetic in effect (one WARNING, W taint) plus the defeated worker
disable. Teardown continued and completed; nothing hung.
== 3. Three observations from reviewing v2 against the 7.0.14 sources ==
(1) mt7925_mac_reset_work() has NO MT76_REMOVED check in 7.0.14. The v2
description says the new mt7921 check aligns mt7921 with mt7925, but on
the mt7925 side the check does not exist, so mt7925 USB keeps the race
that v2 fixes for mt7921. (grep confirms: MT76_REMOVED appears in
mt7925/pci.c and mt792x_dma.c only, not in mt7925/mac.c.)
(2) Cosmetic: mt7921_mcu_parse_response() (mt7921/mcu.c:26) prints
"Message %08x (seq %d) timeout" and calls mt792x_reset() without checking
MT76_MCU_RESET. Because v2 sets MT76_MCU_RESET at disconnect entry,
mt76_mcu_get_response() returns immediately, so every teardown MCU command
logs a timeout ~30 ms after "deregistering interface driver" - in our logs
exactly two per unload: MCU_UNI_CMD(BSS_INFO_UPDATE) (0x00020002) from
interface removal and MCU_EXT_CMD(MAC_INIT_CTRL) (0x000046ed) from
mt792x_stop() -> mt76_connac_mcu_set_mac_enable(). A MT76_MCU_RESET check
before dev_err()/mt792x_reset() would suppress noise that is expected by
design. mt7925/mcu.c:22 has the same print.
(3) Behaviour note: with MT76_REMOVED set on entry, the teardown MCU commands
and mt792xu_wfsys_reset() inside mt792xu_cleanup() fail with -EIO, so a
plain rmmod of a healthy adapter leaves the firmware running. In practice
this is harmless because mt7921u_probe() resets WFSYS when
MT_TOP_MISC2_FW_N9_RDY is set, but it is a behaviour change worth knowing.
== 4. Reproduction attempts: the branch cannot be reached on healthy
hardware ==
We tried to reproduce the WARNING deliberately, to confirm the explanation
above rather than rely on a single capture. Three attempts on 2026-09-27,
all with the v2 modules loaded (mt792x_usb srcversion E7EEAE2D193E09CC1645423):
# method TX load before trigger result
1 sysfs driver unbind 6336 pkt in 10 s (633/s) not
reproduced
(echo 1-9:1.3 > /sys/bus/usb/drivers/mt7921u/unbind)
2 sysfs driver unbind, longer load 7768 pkt in 30 s (258/s) not
reproduced
3 port deauthorisation 6287 pkt in 10 s (628/s) not
reproduced
(echo 0 > /sys/bus/usb/devices/1-9/authorized)
In all three runs the counters were identical:
timed out waiting for pending tx 0
WARNING 0
kthread_park 0
blocked for 0
TX always drained well inside the HZ/5 window, so mt76u_stop_tx() took the
normal path and never called mt76_worker_disable() a second time. The only
kernel output in each run was the expected MCU teardown noise described in
section 3(2): 15 "Message ... timeout" lines per unload, for commands
0x00020002, 0x00020003, 0x00020006 and 0x000046ed. Teardown and re-probe
completed every time; the interface came back after 3 s in all three runs.
This is consistent with the analysis: the .cold branch only executes when TX
cannot drain, i.e. when the chip is dead or stalled, and that state cannot be
induced on demand on healthy hardware. Forcing a disconnect - whether by
unbinding the driver or by deauthorising the port - still leaves the device
answering on the bus, so the queues empty immediately.
In other words, the failed reproduction strengthens rather than weakens the
finding: it is direct evidence of why routine testing (our own 2026-09-24
runs included: three rmmod/insmod cycles plus unplug-during-transfer, no
WARNING at all) cannot surface this path. The 2026-09-26 trace in section 2
remains the only capture of it, obtained during a genuine adapter failure.
== Summary ==
v2 solves the problem it targets: a real adapter death no longer takes the
networking stack or shutdown down with it. The only item we would ask you to
change is the redundant mt76_worker_disable(&dev->mt76.tx_worker) in
mt792xu_disconnect described in section 2.
As a testing hint: the .cold branch can be exercised artificially by
temporarily shortening the HZ/5 timeout in mt76u_stop_tx in a local build,
which may help validate a revised patch without waiting for a hardware
failure.
Happy to test a revised patch on this hardware.
--
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2167595
Title:
mt7921u: USB reset after -110 timeouts deadlocks in mt7921_abort_roc,
blocks rtnl_lock and shutdown
To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2167595/+subscriptions
--
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs