https://bugs.documentfoundation.org/show_bug.cgi?id=173760

            Bug ID: 173760
           Summary: soffice.bin deadlocks in
                    DeInitVCL()/ImplGetSystemDependentDataManager() when
                    closing last document, leaving a zombie instance that
                    blocks all future document opens via IPC
           Product: LibreOffice
           Version: 25.2.7.2 release
          Hardware: All
                OS: Linux (All)
            Status: UNCONFIRMED
          Severity: normal
          Priority: medium
         Component: graphics stack
          Assignee: [email protected]
          Reporter: [email protected]

Summary:
soffice.bin main thread deadlocks in
DeInitVCL()/ImplGetSystemDependentDataManager() when closing the last document,
leaving a permanently running "zombie" instance that blocks all future document
opens via the SingleOfficeIPC pipe

Component: (unsure of exact taxonomy — likely Writer or a general
VCL/graphics-stack component; the defect itself is in shared VCL code, not
Writer-specific)
Version: LibreOffice 25.2.7.2 (Gentoo build,
app-office/libreoffice-25.2.7.2-r1)
OS: Gentoo Linux (OpenRC), KDE Plasma 6 Wayland session, NVIDIA proprietary
driver 595.99.02 (GTX 1650)

Description:

After closing the last open document window in a running LibreOffice session,
the soffice.bin process's main thread can deadlock inside DeInitVCL(), and the
process never actually terminates. It keeps holding the SingleOfficeIPC pipe
and profile lock (~/.config/libreoffice/4/.lock), so it looks like a normal
idle background instance from the outside. Any subsequent attempt to open a
document — for example clicking an attachment in a mail client, which spawns a
fresh oosplash/soffice.bin that hands off the "open this file" request to the
existing instance over the IPC pipe — then hangs forever with only the splash
screen visible, because the instance it's handing off to is permanently wedged
and will never reply.

This is not the first time it has happened; it is a recurring issue, though I
do not yet have a fully deterministic repro. It always presents the same way:
click/open a document while an older LibreOffice instance is idling in the
background, and the new instance hangs at the splash screen indefinitely.

Steps to reproduce (best understanding so far):
1. Open a document in LibreOffice Writer and work in it normally for an
extended period (in the case I captured, the instance had been running for
about 3 days).
2. Close the last open document window (no unusual action taken — just a normal
window close).
3. The soffice.bin process does not exit. It continues to hold the IPC
pipe/profile lock but is actually deadlocked on its main thread.
4. Later, open a different document (e.g. via a file manager, mail client
attachment, or `soffice --writer somefile.odt` from a shell). A new
oosplash/soffice.bin process starts, detects the "running" instance via the IPC
pipe, and hands off the open request — then hangs indefinitely at the splash
screen. No window ever appears.

Actual results:
The new soffice.bin process hangs forever at the splash screen. The old
("master") instance is unkillable via normal means (SIGTERM does nothing useful
since its main thread is not running its normal event loop) — well, it can be
killed with SIGKILL, but that's the only way to clear the situation. The IPC
pipe file and profile lock have to be manually removed as well before a fresh
instance will start cleanly.

Root cause (from live debugging with gdb, see attached backtraces):

The already-running "master" instance's main thread backtrace shows it stuck
inside pthread_mutex_lock(), called from:

  pthread_mutex_lock()
  → basegfx::SystemDependentDataHolder::~SystemDependentDataHolder()
  → SvpSalBitmap::~SvpSalBitmap()
  → Bitmap::~Bitmap()
  → [inlined/unnamed frames in libmergedlo.so]
  → DeInitVCL()
  → ImplSVMain()
  → soffice_main()

I resolved the exact address of the mutex_lock call site against
libmergedlo.so's dynamic symbol table (addr2line / nm), and it lands exactly
in:

  ImplGetSystemDependentDataManager()   (vcl/source/app/svdata.cxx)

which returns a reference to the function-local static
`SystemDependentDataBuffer` singleton that all
`SystemDependentDataHolder`-derived bitmaps register/unregister with via
`startUsage()`/`endUsage()`, each of which locks that object's own `m_aMutex`.

Looking at the current master source for vcl/source/app/svdata.cxx, this exact
mutex is already known to be a source of self-deadlock: `endUsage()`,
`flushAll()`, and the timer handler `implTimeoutHdl()` all contain explicit
workarounds (moving entries out of the map before destructing them, commented
with references to tdf#163428) specifically because destructing a
`SystemDependentData` entry while `m_aMutex` is held can recursively call back
into `endUsage()` on the same mutex and deadlock. Example comment from the
source:

  // we need to destruct the entries outside the lock, because
  // we might call back into endUsage() and that will take the lock again and
deadlock.

The backtrace I captured shows the deadlock happening via a different path than
the ones already hardened (ordinary `Bitmap` destruction during general
`DeInitVCL()` teardown, not via the timer or `flushAll()`), so this may be an
un-hardened variant of the same class of bug, a regression, or simply not yet
backported to 25.2.x. I'm not familiar enough with the surrounding
lifetime/ownership code to say which without guidance from someone who knows
this area.

Meanwhile, the newly-launched soffice.bin's main thread backtrace shows it
blocked exactly where you'd expect for an IPC handoff waiting on a reply that
will never come:

  recv()
  → osl_receivePipe()
  → [frames in libmergedlo.so]
  → InitVCL()
  → ImplSVMain()
  → soffice_main()

Supporting evidence that the master is truly deadlocked (not just slow): its
listening IPC socket showed `Recv-Q: 2` under `ss -xl` — i.e. the kernel had 2
pending, unaccepted connections queued on the socket the master listens on —
while every other thread in the process was cleanly idle (waiting in
poll()/pthread_cond_wait() as expected for a genuinely idle instance). Only the
main thread was wedged, and it happened to be the one responsible for driving
the whole event loop, including the IPC accept/dispatch path.

Full `gdb -p <pid> -batch -ex "thread apply all bt"` output for both processes
is attached (master-bt.txt for the long-running deadlocked instance,
client-bt.txt for the newly launched, blocked-on-handoff instance).

Workaround:
Kill both the old and new `soffice.bin`/`oosplash` processes (SIGKILL is
required for the deadlocked one), then remove the stale IPC pipe and profile
lock files:

  kill -9 <old soffice.bin pid> <old oosplash pid> <new soffice.bin pid> <new
oosplash pid>
  rm -f ~/.config/libreoffice/4/.lock /tmp/OSL_PIPE_<uid>_SingleOfficeIPC_*

After that, a fresh instance starts and opens documents normally, until the
deadlock recurs at some later point.

Additional environment notes:
The deadlocked instance had the full NVIDIA EGL/GLX/OpenCL stack loaded
(libGLX_nvidia, libEGL_nvidia, libnvidia-glcore, libnvidia-opencl, etc.), i.e.
it was using hardware-accelerated rendering rather than a pure software
backend. I don't know whether that's relevant to triggering this particular
mutex's deadlock, but it seemed worth noting given the code in question deals
with GPU/backend-side bitmap caching.

-- 
You are receiving this mail because:
You are the assignee for the bug.

Reply via email to