** Description changed:

  [ Impact ]
  
- QEMU's v8.2.2 coroutine pool implementation can hit the Linux
- vm.max_map_count limit (typically 65560 in Ubuntu), causing QEMU to
- abort with "failed to allocate memory for stack" or "failed to set up
- stack guard page" during coroutine creation. This manifests as
- virtualization hosts not being able to spawn guests.
- 
- This bug has been seen to manifest on large hosts with large guests (32+
- vCPUs).
- 
- The issue seems to be that coroutines can be created but not reused as
- intended. Each coroutine calls mmap(), creating corresponding memory
- mappings that don't go away when the coroutine is no longer needed. The
- coroutine is remaining 'pooled' in a hardware thread by design for later
- reuse (better performance). The problem with this implentation is that
- some threads don't need to keep the coroutines pooled in the first
- place. This effectively 'leaks' the mappings as they cannot be reused
- across thread boundaries in the current implementation.
- 
- The fix upstream switches to a new coroutine pool implementation with a
- global pool that grows to a maximum number of coroutines and per-thread
- local pools that are capped at a hardcoded small number of coroutines.
- Threads that don't need the coroutines 'give them back' to the global
- pool, whereas threads that need them can take them from the global pool.
- This new implementation promotes better reuse by enabling threads to
- return the unneeded coroutines when they are no longer needed, allowing
- the underlying memory mappings to be reused more effectively and
- preventing the need to create more mappings, ultimately preventing the
- exhaustion of the vm.max_map_count limit.
+ QEMU 8.2.2's coroutine pool implementation can hit the Linux
+ vm.max_map_count limit (`65530` in Jammy, `1048576` in Noble), causing
+ QEMU to abort with "failed to allocate memory for stack" or "failed to
+ set up stack guard page" during coroutine creation.
+ 
+ Each coroutine allocates two VMAs with mmap(2). When a coroutine
+ terminates it is cached in a global "release pool"; once that is full,
+ terminated coroutines are cached in the terminating thread's "alloc
+ pool" (up to the max pool size, per thread). When the thread-local pool
+ is empty, the global pool is moved to the thread-local pool. Pooled
+ coroutines are only freed once both pools are full.
+ 
+ The theoretical maximum of coroutines in each thread-local pool is given
+ by: (see hw/block/virtio-blk.c:1638)
+ 
+ 64 + Σ virtio-blk devices (num_queues_i × queue_size_i / 2)
+ 
+ Defaults:
+ 
+ queue_size_i = 256 (hw/block/virtio-blk.c:1722)
+ num_queues_i = vCPUs (hw/virtio/virtio-pci.c:2466)
+ 
+ Assume the virtio-blk defaults for a VM with one disk and 32 vCPUs and
+ you get 4160 coroutines _per IOthread_, plus 8320 for the global pool
+ (2x the max size for a thread-local pool). Assuming the same number of
+ IOThreads as vCPUs (x32) gives a theoretical maximum of 141440 cached
+ coroutines. x2 for two VMAs per coroutine for a grand theoretical max of
+ 282880 VMAs.
+ 
+ In practice the number of coroutines that will actually be allocated is
+ much lower than this, since in order to reach this maximum all of them
+ would need to be active at once. Many small IOs all at once drain both
+ pools, causing more coroutines to be allocated.
+ 
+ This behavior was introduced in upstream 4c41c69e05fe28c (v7.0.0-rc0)
+ and therefore only affects Ubuntu 24.04.
+ 
+ The user who reported this experienced crashes while running Noble's
+ QEMU in a container on a 5.15 kernel (with the old
+ `vm.max_map_count=65530`). Even with a relatively modest VM
+ configuration, the theoretical max of coroutine VMAs is substantially
+ higher than this limit.
+ 
+ Crashes are expected to be significantly less likely with a Noble kernel
+ since reaching the default max_map_count would require the pool's
+ theoretical max size to be increased (using 128+ IOThreads or adding
+ additional virtio-blk devices). Larger VMs might blow past the higher
+ default in Noble's kernel.
+ 
+ While the issue can be worked around by adjusting the sysctl, this is
+ still a resource leak that leaves significant numbers of unused
+ coroutines allocated but unusable. Enforcing that the limit of
+ coroutines be tied to the `vm.max_map_count` hardens VMs against
+ "bursty" IO causing unexpected crashes.
  
  [ Test Plan ]
  
  ```
  lxc launch ubuntu:noble --vm -c limits.cpu=8 -c limits.memory=16GiB n0
  ```
  
  In the VM:
  ```sh
  sudo apt install libvirt-daemon-system virtinst
- 
- # Reduce the max map count to an artificially low number to make reproduction
- # possible on modest hardware
- sudo sysctl -w vm.max_map_count=8192
- sudo sysctl -w vm.max_map_count=24576
  
  wget 
https://cloud-images.ubuntu.com/daily/server/noble/current/noble-server-cloudimg-amd64.img
  sudo cp noble-server-cloudimg-amd64.img 
/var/lib/libvirt/images/testvm-root.qcow2
  sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data.raw 10G
+ sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data2.raw 10G
+ sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data3.raw 10G
+ sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data4.raw 10G
  
  cat > user-data <<EOF
  #cloud-config
  password: ubuntu
  chpasswd:
    expire: false
  ssh_pwauth: true
  package_update: true
  packages:
    - fio
  EOF
  
  sudo virt-install \
    -n testvm \
    --os-variant=ubuntu24.04 \
-   --ram=4096 --vcpus=8 \
-   --iothreads 4 \
+   --ram=8192 --vcpus=8 \
+   --iothreads 8 \
    --import \
    --disk 
path=/var/lib/libvirt/images/testvm-root.qcow2,bus=virtio,cache=writethrough \
    --disk 
path=/var/lib/libvirt/images/testvm-data.raw,bus=virtio,cache=none,driver.io=io_uring
 \
+   --disk 
path=/var/lib/libvirt/images/testvm-data2.raw,bus=virtio,cache=none,driver.io=io_uring
 \
+   --disk 
path=/var/lib/libvirt/images/testvm-data3.raw,bus=virtio,cache=none,driver.io=io_uring
 \
+   --disk 
path=/var/lib/libvirt/images/testvm-data4.raw,bus=virtio,cache=none,driver.io=io_uring
 \
    --graphics none \
    --network network=default \
    --cloud-init user-data=user-data
  
  # Ctrl+Shift+a+]
  
  sudo virsh shutdown testvm
- 
- # Limit writes in the block layer so that writes allocate a coroutine but
- # aren't able to complete right away
- sudo virsh blkdeviotune testvm vdb --total-iops-sec 2000 --current
- ```
- 
- Use `sudo virsh edit testvm` to add the following parameters to the `<driver>`
- section of the 'testvm-data' disk: `queues='8' queue_size='1024' 
iothread='1'`.
- 
- ```sh
+ ```
+ 
+ Limit write throughput in the block layer so that writes allocate a coroutine
+ but aren't able to complete right away:
+ ```sh
+ for disk in vdb vdc vdd vde; do
+   sudo virsh blkdeviotune testvm "${disk}" --total-iops-sec 2000 --current
+ done
+ ```
+ 
+ Use `sudo virsh edit testvm` to pin each disk to its own iothread; otherwise
+ all IO will occur on the main thread. Add `iothread='1'` to the `<driver>`
+ section of each 'testvm-data' disk, incrementing the number for each disk.
+ 
+ ```sh
+ # Reduce the max map count by a factor of four to match the reduction in
+ # used IOThreads in the test environment
+ sudo sysctl -w vm.max_map_count=16384
+ 
  sudo virsh start testvm
  sudo virsh console testvm
  ```
  
  On the host, you can monitor the maps with:
  ```sh
  pid=$(pidof qemu-system-x86_64)
- while sleep 1; do sudo grep -c '' /proc/$pid/maps; done
- ```
- 
- In the testvm, run this fio a few times:
- ```sh
- sudo fio --name=job1 \
-   --ioengine=libaio \
-   --direct=1 \
-   --runtime=15 \
-   --time_based=1 \
-   --rw=randread \
-   --bs=128k \
-   --iodepth=128 \
-   --numjobs=32 \
-   --filename=/dev/vdb
- ```
- 
- Expected behavior:
+ while sleep 1; do sudo grep -c '' "/proc/${pid}/maps"; done
+ ```
+ 
+ In the testvm:
+ ```sh
+ while true; do
+   for dev in vdb vdc vdd vde; do
+     sudo fio "--name=/dev/${dev}" \
+       --ioengine=io_uring \
+       --direct=1 \
+       --runtime=10 \
+       --time_based=1 \
+       --rw=randread \
+       --bs=4k \
+       --iodepth=128 \
+       --numjobs=32 \
+       --group_reporting \
+       "--filename=/dev/${dev}"
+   done
+ done
+ ```
+ 
+ Patched behavior:
  
  /proc/pid/maps levels off and stays consistent through the fio run. (With 
these
- parameters we observed 8685).
- 
- Actual behavior:
+ parameters we observed approximately 5376).
+ 
+ Unpatched behavior:
  
  /proc/pid/maps increases with each fio run until the VM crashes with one
  of:
  
  - failed to allocate memory for stack: Cannot allocate memory
  - failed to set up stack guard page: Cannot allocate memory
  
+ Performance regression test (results should be provided both before & after):
+ ```sh
+ for bs in 4k 128k; do
+   for iodepth in 1 8 16 32 64; do
+     sudo fio --output-format=json --output=fio-$bs-$iodepth.json \
+       --ioengine=libaio --direct=1 --numjobs=$(nproc) \
+       --runtime=30 --ramp_time=5 --time_based=1 \
+       --rw=randread --bs=$bs --iodepth=$iodepth \
+       --name=job1 --filename=/dev/vdb
+   done
+ done
+ ```
+ 
  [ Where problems could occur ]
  
  The patches restructure the way coroutines are allocated and deallocated. 
Coroutines are used extensively in the QMP/HMP protocol implementation, block 
drivers and live migrations; this change is somewhat fundamental to QEMU's 
coroutine infrastructure. If the change were wrong/broken, it could cause:
- - failing block device drivers, casuing failed boots or broken I/O operations 
within affected guests
+ - failing block device drivers, causing failed boots or broken I/O operations 
within affected guests
  - incorrect/strange behavior when interacting with a QEMU process over QMP/HMP
  - failed/broken live migrations
  - performance regressions as the patch is almost entirely changes in the 
coroutine pool hot path (coroutine_pool_get)
  
  In addition to the above hypotheticals, this change will have the
  concrete impact of potentially reducing performance for workloads which
  require many coroutines (as the patches set a hard limit on the size of
- the coroutine pool, so corountines needed above the pool size are
+ the coroutine pool, so coroutines needed above the pool size are
  allocated/deallocated on demand). The coroutine pool size is a function
  of `vm.max_map_count`, so increasing the max map count is the
- reccommended workaround:
+ recommended workaround:
  
  ```
  sysctl -w vm.max_map_count=1048576
  ```
  
  It's likely that workloads running into this performance regression
  would also have been affected by the original bug.
  
- Upstream has two follow-ups to to the refactor,
+ Upstream has two follow-ups to the refactor,
  9352f80cd926fe2dde7c89b93ee33bb0356ff40e and
  25bc7d16fa96b0ff881c83ed225ea380fe427c78. 9352f80cd will be included in
  the upload; 25bc7d16fa fixes a false-positive compiler warning.
  
  [ Other Info ]
  
  QEMU's coroutines can be considered akin to a userspace threading
  implementation. They are used frequently in QEMU's virtual disk drivers.
+ 
+ Noble's QEMU 8.2.2 keeps a thread-local list of coroutines `alloc_pool`
+ (util/qemu-coroutine.c:40); this list can grow to thousands or even tens
+ of thousands of coroutines _for each thread_. Each coroutine has two
+ VMAs created with mmap(2). If there are enough threads available to the
+ QEMU process, the number of coroutines may cause the total number of
+ mappings to exceed `vm.max_map_count`, causing the QEMU process to fail
+ to allocate memory and crash.
  
  From upstream, this bug occurs because "per-thread pools can grow to
  tens of thousands of coroutines. Each coroutine causes 2 virtual memory
  areas to be created. Eventually vm.max_map_count is reached and memory-
  related syscalls fail. The per-thread pool sizes are non-uniform and
  depend on past coroutine usage in each thread, so it's possible for one
  thread to have a large pool while another thread's pool is empty."
  
  The new approach "does not leave large numbers of coroutines pooled in a
  thread that may not use them again. In order to perform well it
  amortizes the cost of global pool accesses by working in batches of
  coroutines instead of individual coroutines."
  
  Upstream commit (coroutine: cap per-thread local pool size): 
https://gitlab.com/qemu-project/qemu/-/commit/86a637e4
  Upstream commit (coroutine: reserve 5,000 mappings): 
https://gitlab.com/qemu-project/qemu/-/commit/9352f80c
+ Upstream commit (util/coroutine: fix -Werror=maybe-uninitialized 
false-positive): https://gitlab.com/qemu-project/qemu/-/commit/25bc7d

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2168803

Title:
  QEMU 8.2.2 fails to start guests by exhausting vm.max_map_count with
  coroutines

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/qemu/+bug/2168803/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to