https://bugs.kde.org/show_bug.cgi?id=525025

            Bug ID: 525025
           Summary: kdeconnectd deadlocks inside the kdeconnect_sms plugin
                    after a bulk conversation query, keeps its bus name
                    but answers nothing
    Classification: Applications
           Product: kdeconnect
      Version First 26.08.0
       Reported In:
          Platform: Other
                OS: Linux
            Status: REPORTED
          Severity: major
          Priority: NOR
         Component: common
          Assignee: [email protected]
          Reporter: [email protected]
                CC: [email protected]
  Target Milestone: ---

kdeconnectd stops answering every D-Bus call while the process stays alive and
keeps its
name on the session bus. This is not a crash. The PID does not change, nothing
appears in
the journal, the process sits at 0 percent CPU, and the phone is still
reachable on the LAN
the whole time. Every call times out with org.freedesktop.DBus.Error.NoReply,
including
kdeconnect-cli --list-devices from a completely unrelated process. The only way
out is to
kill the daemon and let D-Bus activation start a fresh one.

It happened twice in one day here, in two unrelated client sessions. Both times
the last
thing done through the conversations interface was a deep search, meaning
requestConversation
issued for a large number of distinct threads while the thread heads from
requestAllConversationThreads were still streaming in from the phone. The
second time the
search call itself returned normally and the daemon only froze afterwards,
while nothing was
touching it at all, so the connection to the trigger is easy to miss entirely.

A backtrace of the live deadlock shows a circular wait contained entirely in
kdeconnect_sms.so:

  Thread 1 "kdeconnectd":
    QBasicReadWriteLock::contendedTryLockForWrite (libQt6Core)
    ??? (kdeconnect_sms.so)
    Device::privateReceivedPacket(NetworkPacket const&)
(libkdeconnectcore.so.26)
    LanDeviceLink::dataReceived()
    ... QCoreApplication::exec()

  Threads 3, 4, 5 "QThread":
    QWaitCondition::wait(QReadWriteLock*, QDeadlineTimer) (libQt6Core)
    ??? (kdeconnect_sms.so)
    ??? (kdeconnect_sms.so)

  Thread 2 "QThread":
    QLatch::waitInternal(int) (libQt6Core)
    ??? (libQt6DBus)
    QDBusConnection::asyncCall(QDBusMessage const&, int) const
    QDBusAbstractInterface::asyncCallWithArgumentList(...)
    ??? (kdeconnect_sms.so)

Read together: the main thread has just received a network packet from the
phone and wants
the write lock held on the SMS plugin side. The worker threads sit on the read
side waiting
for a condition. One of those workers is blocked inside
QDBusConnection::asyncCall on a
QLatch, and that call can only complete once the main thread services its event
loop, which
it cannot do because it is waiting for the write lock. Nothing can make
progress.

Note that this is a different problem from the file descriptor exhaustion I
filed as bug
525021. The daemon had only 35 descriptors open when it froze and there was no
load on it
at the time, so the two are unrelated apart from both being reachable through
the
conversations API.

Steps to reproduce:

I do not have a deterministic recipe, which is why this report leans on the
backtrace. What
preceded both occurrences was the same:

1. call requestAllConversationThreads() on a device with many SMS threads (1731
here)
2. while the heads are still arriving, issue requestConversation(tid, 0, 200)
for a large
   number of distinct threads, a few in flight at a time
3. let the client finish and go idle

The freeze appeared within the hour, once during the calls and once well after
they had
returned successfully.

Expected result:

The daemon keeps answering D-Bus calls. A conversation query that arrives while
the plugin
is busy is queued or rejected, and no ordering of packets from the phone
against client
requests can leave the plugin lock held.

Actual result:

Permanent deadlock. The daemon holds its bus name and looks healthy to anything
that only
checks whether the name is present, so clients see nothing but timeouts and
users see the
phone silently stop working.

The full backtrace of all ten threads is attached. It was taken without debug
symbols, so
the frames inside kdeconnect_sms.so are unresolved. I can repeat the capture
with debuginfo
installed if that would help.

Environment:

openSUSE Tumbleweed, kernel 7.2.0, Wayland session
kdeconnect-kde 26.08.0
Plasma 6.7.4, Qt 6.11.2, KF6 6.29.0
Phone: Volla Phone X23, Android, 1731 SMS threads

-- 
You are receiving this mail because:
You are watching all bug changes.

Reply via email to