**Complete suspend/resume failure lifecycle on GTX 1060 (Pascal),
580.178.04 — Wayland sessions, kernel 6.8.0-139, systematically
instrumented**
Hardware: GTX 1060 6 GB desktop (ASRock Z270 Pro4), HP E22 G4 on
DisplayPort.
Software:
- Ubuntu 24.04 (Noble), kernel 6.8.0-139-generic
- Driver 580.178.04, stock config:
`NVreg_PreserveVideoMemoryAllocations=1`, `TemporaryFilePath=/var/tmp`
- GNOME on Wayland for all user sessions
(verified live via `XDG_SESSION_TYPE=wayland`; the fossil
`~/.local/share/xorg/Xorg.1.log` shows one Xorg user-server from
the Sep-21 morning, ending `Server terminated successfully` —
no user-session X server has started since).
Side note for GDM: the *greeter* has been X11-stuck since Sep 21
(an early Wayland-greeter failure made GDM fall back; no
`gdm-wayland-session` in any boot since), yet it keeps delivering
Wayland user sessions without issue.
Regression bracket (from dpkg.log):
- 580.173.02 → 580.178.04 landed 2026-08-22 20:10.
- Failure catalog opens ~Sep 11, ~3 weeks later.
- No driver or kernel change in between.
- This independently matches a regression another Pascal reporter
already bracketed to the same release (NVIDIA forum 380301, below).
Method:
- A post-resume hook (drop-in on `nvidia-resume.service`) records NVML
display state and counts NVRM Xids after every resume, saving evidence
snapshots on anomaly.
- A `OnFailure=` trap on `systemd-suspend.service` records failed suspends.
- Panic-time kernel dumps go to efi_pstore, archived under
`/var/lib/systemd/pstore/<epoch>/`
(validated with a deliberate sysrq panic).
The failure is a race: clean and failed cycles interleave at random.
Suspend duration is irrelevant (clean cycles of 4 h and 20 min; failures
after 20 min–11.5 h sleeps). Recently ~1 failure per 2–4 cycles.
Four observed signatures, in escalation order:
1. **Resume renders black screen (compositor and kernel fine).**
Session resumes, compositor alive, but the DisplayPort link is never
retrained; NVML says `Display Active: Disabled` while the compositor
keeps rendering. A monitor power-cycle (HPD) retrains the link and
recovers — repeatedly.
(This same signature was just reported in this very bug on 580.126.09:
black screen every ~3–4 resumes, box reachable over SSH,
display-manager restart restores the picture.)
2. **Framebuffer corruption ("psychedelic colours") + channel poison.**
During the `PreserveVideoMemoryAllocations=0` experiment, on resume:
`NVRM: Xid 13` storm (≈130 events,
`Graphics Exception: SAVE_RESTORE_ADDR_OOB`, pid=gnome-shell).
Screen shows garbage; GNOME survived a `systemctl restart gdm`.
3. **Suspend freeze-abort with the compositor as the refusing task.**
Recurrent (4+ times):
`Freezing user space processes failed after 20.001 seconds
(2 tasks refusing to freeze)`
with `task:gnome-shell state:R` spinning inside NVIDIA RM paths
(`os_acquire_spinlock` / `_nv*rm [nvidia]` in the stack),
plus `NVRM: Xid 158: timeout error waiting for
NV_UFLUSH_FB_FLUSH` ×5–8 from gnome-shell.
`systemd-suspend.service` spends ~8 minutes failing to SIGKILL its
`nvidia-sleep.sh` child (uninterruptible), then fails;
`nvidia-resume.service` hangs too.
The box is left awake, wedged, fans spinning; only a hardware
power-off exits.
Notably, a queued unit in that transaction even made systemd
**refuse a Ctrl+Alt+Del reboot**
("Transaction for reboot.target/start is destructive").
4. **Poisoned wake after a successful s2idle sleep.**
The automatic s2idle retry *did* sleep (16 min, journal silent),
but the wake killed the display: `Xid 158` burst +
`[drm:__nv_drm_gem_nvkms_map] Failed to map NvKms...`,
compositor died mid-greeter-respawn. Power-off.
Corroborating pattern: a resume can also silently kill an *application*
channel — `Xid 69, pid=X, name=Discord, Class Error: channel 0x38,
ErrorCode 4` seconds after returning from a clean 4-hour sleep, display
perfectly healthy (NVML `Enabled`) — and the *next* suspend then fails
with signature 3. Both post-resume Xid variants (69 from an app, 158 from
the compositor) preceded suspend freeze-aborts in every observed case.
This is **not** the VRAM-restore race alone: setting
`PreserveVideoMemoryAllocations=0` changed the symptoms (Xid 13 storm
instead of black screen) but not the disease. The failure reproduces with
stock settings; the driver-side channel state is what wedges — the
freezer splat is a symptom, not the cause.
Independent corroboration — same driver, same GPU generation,
same signatures:
- NVIDIA forum 380301: GTX 1060 Mobile (Pascal), Hyprland/Wayland,
580.178.04 — `Xid 62 (PMU halt)` → repeating
`Xid 158 NV_UFLUSH_FB_FLUSH` → Xid 154, frozen screen, reboot-only;
first occurrence on the boot right after 580.159.04 → 580.178.04.
- NVIDIA forum 382518: GTX 1070 (Pascal), Plasma 6/Wayland, 580.178.04 —
`nvidia-drm: Failed to allocate NVKMS memory for GEM object` under
compositor load → Wayland freezes, plasmashell crash
(same NVKMS/GEM failure as signature 4's
`[drm:__nv_drm_gem_nvkms_map] Failed to map NvKmsKapiMemory`).
- NVIDIA open-gpu-kernel-modules #1137: suspend/resume freeze on
Wayland+GNOME, open-driver branch — same family, driver-side.
Three compositors (GNOME, Plasma, Hyprland) × Pascal × 580.178.04 →
one driver-side cause.
--
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2139645
Title:
Black screen on resume since NVIDIA 580.126.09 update
To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/nvidia-graphics-drivers-570/+bug/2139645/+subscriptions
--
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs