Hi,

On an ASUS Zenbook A16 (UX3607OA, Snapdragon X2 Elite Extreme "Glymur", Adreno 
GPU, chip id 0x44070041) we are hitting an unrecoverable GPU recovery loop 
that, once triggered, can only be cleared by a full reboot.

Environment:
- kernel: linux-next next-20260730, plus a handful of local patches for this 
laptop (USB wakeup-source/dma-coherent DTS, CPU thermal trips, an eDP 
link-training retry tweak and a DP debug-print ratelimit); none of them touch 
the GPU, GMU, GDSC or GPU clock paths
- Mesa 26.1.5 (freedreno, a8xx)
- GMU firmware v5.2.38 (qcom/gen80100_gmu.bin) and gen80100_sqe.fw from 
linux-firmware
- niri (smithay-based Wayland compositor), ordinary desktop clients

Symptom: a regular client (rotating across sessions: chromium, firefox, 
Xwayland, thunderbird) eventually triggers a GPU fault. From that point 
recovery loops at roughly 1 Hz; every iteration logs:

  adreno 3d00000.gpu: [drm:a8xx_irq [msm]] *ERROR* gpu fault ring 2 fence 28f9c 
status 00880005 gfx_status 00880004
  adreno 3d00000.gpu: [drm:a8xx_irq [msm]] *ERROR* BR: status 00880004 rb 
0002/0078 ib1 0000000000000000/0000 ib2 0000000000000000/0000 ib3 
0000000000000000/0000
  adreno 3d00000.gpu: [drm:a8xx_irq [msm]] *ERROR* BV: status 00880004 rb 
0000/0078 ib1 0000000000000000/0000 ib2 0000000000000000/0000 ib3 
0000000000000000/0000
  msm_dpu ae01000.display-controller: [drm:recover_worker [msm]] *ERROR* 
44070041: hangcheck recover!
  msm_dpu ae01000.display-controller: [drm:recover_worker [msm]] *ERROR* 
44070041: offending task: <client>
  adreno 3d00000.gpu: [drm:a8xx_recover [msm]] *ERROR* cx gdsc didn't collapse

Observations across four separate incidents (2026-08-09 to 2026-08-20; longest 
60+ minutes; worst measured rate 161 recover cycles in 5 minutes):

1. The reported offending client differs between incidents and can change 
within a single incident.
2. Killing the offending client does not stop the loop. We verified this 
explicitly: after SIGTERM-ing the reported victim the loop continued at an 
unchanged rate within seconds, and the next client that submitted work became 
the new "offending task". The corrupted state evidently survives in 
hardware/firmware rather than in the client.
3. The loop never converges on its own; only a full reboot restores the GPU. 
While it runs, the victim's surface stays black and the whole desktop degrades, 
since every recovery stalls the GPU.
4. Onset correlates with many clients creating GPU contexts in parallel, e.g. 
right after a session restore spawns half a dozen applications at once.

Since every iteration reports "cx gdsc didn't collapse", the recovery path 
apparently never succeeds in power-cycling the GPU, which would explain why the 
fault state persists across recovery attempts. Two questions this raises:

- Is the CX GDSC on Glymur shared with, or kept up by, other consumers (the 
display path, for instance) so that it cannot collapse while the panel is 
active - which would make this recovery flow structurally ineffective on this 
SoC?
- Or does gen8/GMU require a different stop/reset sequence before the collapse 
is attempted?

We are happy to test patches, run instrumented builds, and provide full 
journals or devcoredumps - the machine reproduces this every couple of days 
under normal desktop use.

Thanks,
Greg Ociepka

Reply via email to