https://bugs.kde.org/show_bug.cgi?id=525025

--- Comment #2 from Branislav Klocok <[email protected]> ---
Follow-up with a third occurrence, this one provoked on purpose, and with what
looks like the
knob that controls it.

I reproduced the freeze deliberately today with the same client, and the number
of concurrent
requestConversation calls made the difference. With six requests in flight over
200 distinct
threads the daemon froze exactly as before. With two in flight over the same
200 threads it
survived two consecutive runs, 68 and 57 seconds, and stayed responsive
afterwards. That is one
failure against two successes, so it is evidence rather than proof, but it is
the first handle I
have found on the timing.

The new backtrace is attached and matches the earlier one in shape, with one
detail that stands
out. Where the first capture had three worker threads waiting on
QWaitCondition(QReadWriteLock*),
this one has nine, plus a tenth blocked in QDBusConnection::asyncCall on the
QLatch, and the main
thread in Device::privateReceivedPacket wanting the write lock as before. The
LWP ids show those
workers were created in batches during the run. The plugin therefore appears to
spawn a worker
thread per conversation request and not to reap them while they are blocked, so
the number of
threads piled on the lock scales with how many requests are in flight. That
fits the concurrency
dependence above.

Conditions matched the earlier two occurrences: requestAllConversationThreads
was still streaming
thread heads from the phone while the requestConversation calls were going out,
so packets kept
arriving into privateReceivedPacket throughout.

For anyone hitting this before it is fixed, keeping the number of conversation
requests in flight
low is what worked here. I have set my own client to two, and I now probe the
daemon for liveness
after every bulk query, because a deadlocked daemon is otherwise
indistinguishable from an empty
mailbox and a client will happily report "no messages found".

-- 
You are receiving this mail because:
You are watching all bug changes.

Reply via email to