** Description changed:

  [ Impact ]
  
-  * An explanation of the effects of the bug on users and justification
-    for backporting the fix to the stable release.
+ QEMU's v8.2.2 coroutine pool implementation can hit the Linux
+ vm.max_map_count limit (typically 65560 in Ubuntu), causing QEMU to
+ abort with "failed to allocate memory for stack" or "failed to set up
+ stack guard page" during coroutine creation. This manifests as
+ virtualization hosts not being able to spawn guests.
  
-  * In addition, it is helpful, but not required, to include an
-    explanation of how the upload fixes this bug.
+ This bug has been seen to manifest on large hosts with large guests (32+
+ vCPUs).
  
- [ Test Plan ]
+ The issue seems to be that coroutines can be created but not reused as
+ intended. Each coroutine calls mmap(), creating corresponding memory
+ mappings that don't go away when the coroutine is no longer needed. The
+ coroutine is remaining 'pooled' in a hardware thread by design for later
+ reuse (better performance). The problem with this implentation is that
+ some threads don't need to keep the coroutines pooled in the first
+ place. This effectively 'leaks' the mappings as they cannot be reused
+ across thread boundaries in the current implementation.
  
-  * detailed instructions how to reproduce the bug
+ The fix upstream switches to a new coroutine pool implementation with a
+ global pool that grows to a maximum number of coroutines and per-thread
+ local pools that are capped at a hardcoded small number of coroutines.
+ Threads that don't need the coroutines 'give them back' to the global
+ pool, whereas threads that need them can take them from the global pool.
+ This new implementation promotes better reuse by enabling threads to
+ return the unneeded coroutines when they are no longer needed, allowing
+ the underlying memory mappings to be reused more effectively and
+ preventing the need to create more mappings, ultimately preventing the
+ exhaustion of the vm.max_map_count limit.
  
-  * these should allow someone who is not familiar with the affected
-    package to reproduce the bug and verify that the updated package
-    fixes the problem.
+ [ Test Plan WIP ]
  
-  * if other testing is appropriate to perform before landing this
-    update, this should also be described here.
+ detailed instructions how to reproduce the bug
  
- [ Where problems could occur ]
+ these should allow someone who is not familiar with the affected package
+ to reproduce the bug and verify that the updated package fixes the
+ problem.
  
-  * Think about what the upload changes in the software. Imagine the
-    change is wrong or breaks something else: how would this show up?
+ if other testing is appropriate to perform before landing this update,
+ this should also be described here.
  
-  * It is assumed that any SRU candidate patch is well-tested before
-    upload and has a low overall risk of regression, but it's important
-    to make the effort to think about what ''could'' happen in the event
-    of a regression.
+ [ Where problems could occur WIP ]
  
-  * This must never be "None" or "Low", or entirely an argument as to why
-    your upload is low risk.
+ TBD; Think about what the upload changes in the software. Imagine the
+ change is wrong or breaks something else: how would this show up?
  
-  * This both shows the SRU team that the risks have been considered,
-    and provides guidance to testers in regression-testing the SRU.
+ It is assumed that any SRU candidate patch is well-tested before upload
+ and has a low overall risk of regression, but it's important to make the
+ effort to think about what ''could'' happen in the event of a
+ regression.
+ 
+ This must never be "None" or "Low", or entirely an argument as to why
+ your upload is low risk.
+ 
+ This both shows the SRU team that the risks have been considered, and
+ provides guidance to testers in regression-testing the SRU.
  
  [ Other Info ]
  
-  * Anything else you think is useful to include
+ QEMU's coroutines can be considered akin to a userspace threading
+ implementation. They are used frequently in QEMU's virtual disk drivers.
  
-  * Make sure to explain any deviation from the norm, to save the SRU
-    reviewer from having to infer your reasoning, possibly incorrectly.
-    This should also help reduce review iterations, particularly when the
-    reason for the deviation is not obvious.
+ From upstream, this bug occurs because "per-thread pools can grow to
+ tens of thousands of coroutines. Each coroutine causes 2 virtual memory
+ areas to be created. Eventually vm.max_map_count is reached and memory-
+ related syscalls fail. The per-thread pool sizes are non-uniform and
+ depend on past coroutine usage in each thread, so it's possible for one
+ thread to have a large pool while another thread's pool is empty."
  
-  * Anticipate questions from users, SRU, +1 maintenance, security teams
-    and the Technical Board and address these questions in advance
+ The new approach "does not leave large numbers of coroutines pooled in a
+ thread that may not use them again. In order to perform well it
+ amortizes the cost of global pool accesses by working in batches of
+ coroutines instead of individual coroutines."
+ 
+ Upstream commit (coroutine: cap per-thread local pool size): 
https://gitlab.com/qemu-project/qemu/-/commit/86a637e4
+ Upstream commit (coroutine: reserve 5,000 mappings): 
https://gitlab.com/qemu-project/qemu/-/commit/9352f80c

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2168803

Title:
  QEMU 8.2.2 crashes by exhausting vm.max_map_count with coroutines

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/qemu/+bug/2168803/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to