https://bugs.kde.org/show_bug.cgi?id=525553

Max Nomad <[email protected]> changed:

           What    |Removed                     |Added
----------------------------------------------------------------------------
                 CC|                            |[email protected]

--- Comment #14 from Max Nomad <[email protected]> ---
Not the reporter, but I hit the same failure and I think my configuration is
useful here precisely because it is the opposite end of the version range from
Matthew's. Posting to help bound the problem rather than to pile on a me-too.

Plasma 6.6.6 / Frameworks 6.24.0 / Qt 6.10.2
Ubuntu 26.04.1 LTS, kernel 7.0.0-34
Wayland session
NVIDIA GA106M [RTX 3060 Mobile] (Ampere), proprietary driver 580.178.04
Hybrid laptop (AMD Cezanne iGPU + NVIDIA dGPU); KWin renders on the dGPU
NVreg_PreserveVideoMemoryAllocations=1, nvidia-suspend/resume enabled
On "other people have seen it resolved after an update to the NVIDIA drivers"
I do not think a driver update is the fix, and the two reports in this bug are
evidence against it:

Matthew (comment 0)     me
GPU     RTX 2080 Ti (Turing, desktop)   RTX 3060 Mobile (Ampere, laptop)
Driver  nvidia-open 610.57.04 (newest branch)   nvidia proprietary 580.178.04
Distro / kernel Fedora 44 / 7.2.4       Ubuntu 26.04 / 7.0.0-34
Plasma / Qt     6.7.5 / 6.11.2  6.6.6 / 6.10.2
Different GPU generation, different driver branch, different kernel, different
distro, and Plasma/Qt one release apart in each direction — with the same
failure signature. If this were a driver regression that an update fixes, one
of these two configurations should be clean. Neither is.

To be explicit about my own limits: I am deliberately not on the latest driver
or the latest Plasma, so I cannot answer "does the newest stack fix it" — but
Matthew already is on the newest stack, and it does not.

Reproduction rate
17 of 21 suspend/resume cycles over 10 days (81%) on my machine. I have a
systemd hook watching the journal, so every occurrence is timestamped. This is
not intermittent here; it is close to every resume.

The backtrace that was requested
Attached: plasmashell-qfatal-backtrace.txt — full thread apply all bt, 15
threads, symbolized with libqt6quick6-dbgsym 6.10.2+dfsg-3 plus debuginfod, so
the Qt Quick frames carry file:line.

I could not get a backtrace of a frozen process in the end, and the reason is
itself a data point: on this machine the failure takes two forms, and the
common one does not leave a frozen process to attach to. It qFatals instead.
The crashing path is:

#10 __GI___abort ()                                          at
stdlib/abort.c:77
#12 QMessageLogger::fatal(char const*, ...) const           (libQt6Core.so.6)
#13 QSGRenderLoop::handleContextCreationFailure(QQuickWindow*)
                                    at qsgrenderloop.cpp:292
#14 QSGGuiThreadRenderLoop::ensureRhi(QQuickWindow*, WindowData&)
                                    at qsgrenderloop.cpp:496
#15 QSGGuiThreadRenderLoop::renderWindow(QQuickWindow*)
                                    at qsgrenderloop.cpp:580
#16 QSGGuiThreadRenderLoop::exposureChanged(QQuickWindow*)
                                    at qsgrenderloop.cpp:779
#17 QWindow::event(QEvent*)                                 (libQt6Gui.so.6)
#20 QGuiApplicationPrivate::processExposeEvent(...)         (libQt6Gui.so.6)
#22 QWindowSystemInterface::handleExposeEvent<SynchronousDelivery>(QWindow*,
QRegion const&)
#23 QtWaylandClient::QWaylandWindow::sendExposeEvent(QRect const&)
#24 QtWaylandClient::QWaylandWindow::updateExposure()
qsgrenderloop.cpp:496 is ensureRhi() failing to create the RHI after the device
loss, which calls handleContextCreationFailure() at :292, which is a qFatal —
hence Failed to create RHI (backend 2) and an immediate abort(). Note this is
QSGGuiThreadRenderLoop, the basic (non-threaded) render loop.

The other form logs Graphics device lost, cleaning up scenegraph and releasing
RHIs and then simply never renders again — that one does sit frozen, and is
what this bug and 525886 describe. 14 of 20 occurrences here took that path, 6
took the qFatal path. Both start from the same QRhiGles2: Context is lost.

Post-mortem traces from the frozen variant, after it is asked to quit, land in:

SIGSEGV in QtQuick/QML item code, consistent with QML bindings still driving
items whose scenegraph was already released:

#4  QQuickItem::setAcceptedMouseButtons(...)                (libQt6Quick.so.6 +
0x214fd4)
#10 QQuickItem::setWidth(double)                            (libQt6Quick.so.6 +
0x217580)
#15 QQmlBinding::doUpdate(...)                              (libQt6Qml.so.6 +
0x2e1d52)
The GPU is not dead when this happens
Worth weighing against closing this as a driver bug: in 10 days and 54 logged
context-loss lines on this machine, every single one came from plasmashell. No
other process logged a lost GL or EGL context.

KWin is rendering on the same RTX 3060, in the same session, across all 21 of
those resumes, and never drops a frame — it takes the device loss and carries
on. There were no Xid errors and no NV_ERR_RESET_REQUIRED on any affected
resume. So the driver does something at resume that invalidates contexts, KWin
recovers from it, and plasmashell does not. That gap looks like it is on the Qt
Quick side.

Timeline of one occurrence, for the ~14s delay:

08:30:48.071  kernel: PM: suspend exit
08:31:02.825  plasmashell: QRhiGles2: Context is lost.
08:31:02.825  plasmashell: Failed to create RHI (backend 2)
The context is not lost at resume — it is lost at plasmashell's first render
attempt afterwards. That gap is consistent across all 17 occurrences.

Happy to provide core dumps, run a patched build, or test against a newer
driver if that would settle it.

Possibly the same underlying issue as bug 525886.

-- 
You are receiving this mail because:
You are watching all bug changes.

Reply via email to