*** This bug is a duplicate of bug 2163303 ***
https://bugs.launchpad.net/bugs/2163303
Public bug reported:
# linux-firmware 2.29: DMCUB firmware fails to load on Navi 21 (RX
6800), causing display corruption and GPU hangs in multi-monitor setups
## Summary
The `sienna_cichlid_dmcub.bin` shipped in `linux-firmware`
`20240318.git3b128b60-0ubuntu2.29` fails to load on Navi 21 (Radeon RX 6800,
DCN 3.0). The kernel reports an invalid DMUB version (`0x01000000`) and
`Wait for DMUB auto-load failed: 3`, then continues operating with the display
microcontroller in a degraded state.
With more than one display active this leads to display plane corruption,
`amdgpu_dm_commit_planes` warnings, MPC register wait timeouts, SMU
communication loss, GPU ring timeouts and eventually a full system freeze.
Downgrading **only** that one firmware file to the version from
`20240318.git3b128b60-0ubuntu2.26` fixes the problem completely.
## Package / versions
- Distro: Linux Mint 22.3 (Ubuntu 24.04 noble base)
- Package: `linux-firmware`
- Bad version: `20240318.git3b128b60-0ubuntu2.29`
- Good version: `20240318.git3b128b60-0ubuntu2.26`
- Affected file: `/lib/firmware/amdgpu/sienna_cichlid_dmcub.bin.zst`
- `.29` md5: `8783824f37745ec5d53ee8a2d71b18f7`
- `.26` md5: `041e8ee4f578b5eccb1bb89a2f14b1db`
Of the 12 `sienna_cichlid_*` firmware files, this is the **only** one that
differs between `.26` and `.29`. All others are byte-identical.
## Hardware
- GPU: `Advanced Micro Devices, Inc. [AMD/ATI] Navi 21 [Radeon RX 6800/6800 XT
/ 6900 XT] [1002:73bf] (rev c3)`
(Sapphire, 16 GB, DCN 3.0, Display Core v3.2.369)
- Motherboard: ASUS ROG STRIX B350-F GAMING, BIOS 5603
- Displays: 3 monitors
- `DP-1` 1920x1080@60 (Philips)
- `DP-2` 1920x1080@60
- `HDMI-A-1` [email protected]
- Session: X11 (Cinnamon)
## Not kernel-dependent
The failure follows the firmware, not the kernel. All three installed kernels
reproduce it while the `.29` firmware is in place:
| Kernel | DMUB version reported | GPU errors in that boot |
|---|---|---|
| 6.14.0-37-generic | `0x01000000` | yes |
| 7.0.0-28-generic | `0x01000000` | yes (47) |
| 7.0.0-29-generic | `0x01000000` | yes (21 / 23 / 24 / 212) |
| 7.0.0-29-generic | `0x02020020` (after firmware downgrade) | **0** |
## Steps to reproduce
1. Navi 21 GPU with `linux-firmware` `2.29` installed.
2. Connect 3 displays (2x DisplayPort, 1x HDMI) and boot into a graphical
session.
3. Use the desktop normally. Anything that triggers a connector re-probe
(for example running `xrandr --query`, or the desktop environment
re-reading the display configuration) can trigger the failure.
## Expected
DMUB firmware loads and reports a valid version; displays work normally.
## Actual
### 1. DMUB fails to load at every boot
```
amdgpu 0000:0b:00.0: [drm] Loading DMUB firmware via PSP: version=0x01000000
amdgpu 0000:0b:00.0: [drm] Wait for DMUB auto-load failed: 3
amdgpu 0000:0b:00.0: [drm] DMUB hardware initialized: version=0x01000000
```
`0x01000000` is not a valid DMUB firmware version.
### 2. Visible symptom
Vertical coloured stripes covering roughly half of one display. Application
windows render *behind* the stripes, while the mouse cursor draws correctly
*in front* of them — consistent with the hardware cursor plane still being
updated while the main plane composition is stuck.
Never reproduced with a single display active (e.g. recovery mode).
Never reproduced under Windows on the same hardware and cables.
### 3. Kernel warning in the plane commit path
```
WARNING: drivers/gpu/drm/amd/amdgpu/../display/amdgpu_dm/amdgpu_dm.c:10181
at amdgpu_dm_commit_planes+0x11c8/0x1740 [amdgpu]
WARNING: drivers/gpu/drm/amd/amdgpu/../display/amdgpu_dm/amdgpu_dm.c:9567
at amdgpu_dm_commit_planes+0x11cf/0x1740 [amdgpu]
Call Trace:
amdgpu_dm_commit_streams+0x658/0x8b0 [amdgpu]
amdgpu_dm_atomic_commit_tail+0xb9/0xda0 [amdgpu]
commit_tail+0xc9/0x1b0
drm_atomic_helper_commit+0x132/0x160
drm_atomic_commit+0xaf/0xf0
dm_restore_drm_connector_state+0x101/0x1b0 [amdgpu]
handle_hpd_irq_helper+0x24e/0x310 [amdgpu]
handle_hpd_irq+0xe/0x20 [amdgpu]
dm_irq_work_func+0x19/0x30 [amdgpu]
process_one_work+0x1af/0x430
worker_thread+0x1bf/0x350
```
### 4. MPC never goes idle
```
amdgpu 0000:0b:00.0: [drm] REG_WAIT timeout 1us * 100000 tries -
mpc2_assert_idle_mpcc line:481
amdgpu 0000:0b:00.0: [drm] REG_WAIT timeout 1us * 10 tries - optc3_lock line:128
workqueue: dm_irq_work_func [amdgpu] hogged CPU for >10000us 11 times
[drm:dcn20_wait_for_blank_complete [amdgpu]] *ERROR* DC: failed to blank crtc!
amdgpu 0000:0b:00.0: [drm] *ERROR* [CRTC:470:crtc-0] flip_done timed out
amdgpu 0000:0b:00.0: [drm] *ERROR* [CRTC:470:crtc-0] commit wait timed out
```
### 5. Cascade into SMU loss and GPU hang
```
amdgpu 0000:0b:00.0: SMU: No response msg_reg: 22 resp_reg: 0 (x210 in one
boot)
amdgpu 0000:0b:00.0: Failed to disable gfxoff!
amdgpu 0000:0b:00.0: Failed to set workload mask 0x00000001
amdgpu 0000:0b:00.0: (-62) failed to disable fullscreen 3D power profile mode
amdgpu 0000:0b:00.0: ring gfx_0.0.0 timeout, signaled seq=25441, emitted
seq=25441
amdgpu 0000:0b:00.0: ring sdma2 timeout, signaled seq=620, emitted seq=622
amdgpu 0000:0b:00.0: [drm:amdgpu_ring_test_helper [amdgpu]] *ERROR* ring
kiq_0.2.1.0 test failed (-110)
amdgpu 0000:0b:00.0: [drm] device wedged, but recovered through reset
```
### 6. PCIe AER errors from the GPU and its audio function
```
pcieport 0000:00:03.1: AER: Multiple Uncorrectable (Non-Fatal) error message
received from 0000:0b:00.1
amdgpu 0000:0b:00.0: PCIe Bus Error: severity=Uncorrectable (Non-Fatal),
type=Transaction Layer, (Requester ID)
amdgpu 0000:0b:00.0: device [1002:73bf] error status/mask=00100000/00000000
snd_hda_intel 0000:0b:00.1: AER: TLP Header: 0x40001001 0x0000000c 0xfca2000c
0x00000000
pcieport 0000:0a:00.0: AER: device recovery failed
```
The system then freezes hard — screens go black, nothing further is written to
the journal, and a forced power-off is required.
## Workarounds that did NOT help
- Booting the previous kernel (`7.0.0-28-generic`) or `6.14.0-37-generic`
- Xorg `Option "EnablePageFlip" "off"` / `"TearFree" "false"` /
`"AsyncFlipSecondaries" "false"`
- `amdgpu.ppfeaturemask` with GFXOFF disabled
- `amdgpu.dcdebugmask=0x10` + `amdgpu.runpm=0` + `amdgpu.gpu_recovery=1`
Note: `amdgpu.dc=0` is **not** a viable workaround on this hardware — Display
Core is mandatory for DCN, and the parameter results in no display output at
all.
## Workaround that DOES work
Replace only `/lib/firmware/amdgpu/sienna_cichlid_dmcub.bin.zst` with the file
from `linux-firmware` `20240318.git3b128b60-0ubuntu2.26`, regenerate the
initramfs and reboot:
```
apt-get download linux-firmware=20240318.git3b128b60-0ubuntu2.26
dpkg-deb -x linux-firmware_*2.26_amd64.deb fw26/
cp fw26/lib/firmware/amdgpu/sienna_cichlid_dmcub.bin.zst \
/lib/firmware/amdgpu/sienna_cichlid_dmcub.bin.zst
update-initramfs -u -k all
```
After this the DMUB loads correctly and all errors disappear:
```
amdgpu 0000:0b:00.0: [drm] Loading DMUB firmware via PSP: version=0x02020020
amdgpu 0000:0b:00.0: [drm] DMUB hardware initialized: version=0x02020020
```
No `Wait for DMUB auto-load failed`, no plane commit warnings, no SMU timeouts,
no freezes, with all 3 displays active.
## Impact
Any Navi 21 (RX 6800 / 6800 XT / 6900 XT) system on Ubuntu 24.04 / derivatives
with `linux-firmware` `2.29` and more than one display is likely affected. The
symptom is a hard system freeze, so it can cause data loss.
## Note on the 2.27 version
`20240318.git3b128b60-0ubuntu2.27` (the version that introduced the update on
this system) is no longer available in the archive; only `.26`, `.29` and the
original `2` are. The regression was therefore bisected to the `.26` → `.29`
range rather than to an exact intermediate version. The `.27` and `.28`
changelog entries contain numerous "DMCUB updates for various AMDGPU ASICs"
commits.
** Affects: linux-firmware (Ubuntu)
Importance: Undecided
Status: New
--
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2164690
Title:
linux-firmware 2.29: DMCUB firmware fails to load on Navi 21 (RX
6800), causing display corruption and GPU hangs in multi-monitor
setups
To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux-firmware/+bug/2164690/+subscriptions
--
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs