** Description changed:

  [ Impact ]
  
  QEMU's v8.2.2 coroutine pool implementation can hit the Linux
  vm.max_map_count limit (typically 65560 in Ubuntu), causing QEMU to
  abort with "failed to allocate memory for stack" or "failed to set up
  stack guard page" during coroutine creation. This manifests as
  virtualization hosts not being able to spawn guests.
  
  This bug has been seen to manifest on large hosts with large guests (32+
  vCPUs).
  
  The issue seems to be that coroutines can be created but not reused as
  intended. Each coroutine calls mmap(), creating corresponding memory
  mappings that don't go away when the coroutine is no longer needed. The
  coroutine is remaining 'pooled' in a hardware thread by design for later
  reuse (better performance). The problem with this implentation is that
  some threads don't need to keep the coroutines pooled in the first
  place. This effectively 'leaks' the mappings as they cannot be reused
  across thread boundaries in the current implementation.
  
  The fix upstream switches to a new coroutine pool implementation with a
  global pool that grows to a maximum number of coroutines and per-thread
  local pools that are capped at a hardcoded small number of coroutines.
  Threads that don't need the coroutines 'give them back' to the global
  pool, whereas threads that need them can take them from the global pool.
  This new implementation promotes better reuse by enabling threads to
  return the unneeded coroutines when they are no longer needed, allowing
  the underlying memory mappings to be reused more effectively and
  preventing the need to create more mappings, ultimately preventing the
  exhaustion of the vm.max_map_count limit.
  
- [ Test Plan WIP ]
+ [ Test Plan ]
  
- detailed instructions how to reproduce the bug
+ ```
+ lxc launch ubuntu:noble --vm -c limits.cpu=8 -c limits.memory=16GiB n0
+ ```
  
- these should allow someone who is not familiar with the affected package
- to reproduce the bug and verify that the updated package fixes the
- problem.
+ In the VM:
+ ```sh
+ sudo apt install libvirt-daemon-system virtinst
  
- if other testing is appropriate to perform before landing this update,
- this should also be described here.
+ # Reduce the max map count to an artificially low number to make reproduction
+ # possible on modest hardware
+ sudo sysctl -w vm.max_map_count=8192
+ sudo sysctl -w vm.max_map_count=24576
  
- [ Where problems could occur WIP ]
+ wget 
https://cloud-images.ubuntu.com/daily/server/noble/current/noble-server-cloudimg-amd64.img
+ sudo cp noble-server-cloudimg-amd64.img 
/var/lib/libvirt/images/testvm-root.qcow2
+ sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data.raw 10G
  
- TBD; Think about what the upload changes in the software. Imagine the
- change is wrong or breaks something else: how would this show up?
+ cat > user-data <<EOF
+ #cloud-config
+ password: ubuntu
+ chpasswd:
+   expire: false
+ ssh_pwauth: true
+ package_update: true
+ packages:
+   - fio
+ EOF
  
- It is assumed that any SRU candidate patch is well-tested before upload
- and has a low overall risk of regression, but it's important to make the
- effort to think about what ''could'' happen in the event of a
- regression.
+ sudo virt-install \
+   -n testvm \
+   --os-variant=ubuntu24.04 \
+   --ram=4096 --vcpus=8 \
+   --iothreads 4 \
+   --import \
+   --disk 
path=/var/lib/libvirt/images/testvm-root.qcow2,bus=virtio,cache=writethrough \
+   --disk 
path=/var/lib/libvirt/images/testvm-data.raw,bus=virtio,cache=none,driver.io=io_uring
 \
+   --graphics none \
+   --network network=default \
+   --cloud-init user-data=user-data
  
- This must never be "None" or "Low", or entirely an argument as to why
- your upload is low risk.
+ # Ctrl+Shift+a+]
  
- This both shows the SRU team that the risks have been considered, and
- provides guidance to testers in regression-testing the SRU.
+ sudo virsh shutdown testvm
+ 
+ # Limit writes in the block layer so that writes allocate a coroutine but
+ # aren't able to complete right away
+ sudo virsh blkdeviotune testvm vdb --total-iops-sec 2000 --current
+ ```
+ 
+ Use `sudo virsh edit testvm` to add the following parameters to the `<driver>`
+ section of the 'testvm-data' disk: `queues='8' queue_size='1024' 
iothread='1'`.
+ 
+ ```sh
+ sudo virsh start testvm
+ sudo virsh console testvm
+ ```
+ 
+ On the host, you can monitor the maps with:
+ ```sh
+ pid=$(pidof qemu-system-x86_64)
+ while sleep 1; do sudo grep -c '' /proc/$pid/maps; done
+ ```
+ 
+ In the testvm, run this fio a few times:
+ ```sh
+ sudo fio --name=job1 \
+   --ioengine=libaio \
+   --direct=1 \
+   --runtime=15 \
+   --time_based=1 \
+   --rw=randread \
+   --bs=128k \
+   --iodepth=128 \
+   --numjobs=32 \
+   --filename=/dev/vdb
+ ```
+ 
+ Expected behavior:
+ 
+ /proc/pid/maps levels off and stays consistent through the fio run. (With 
these
+ parameters we observed 8685).
+ 
+ Actual behavior:
+ 
+ /proc/pid/maps increases with each fio run until the VM crashes with one
+ of:
+ 
+ - failed to allocate memory for stack: Cannot allocate memory
+ - failed to set up stack guard page: Cannot allocate memory
+ 
+ [ Where problems could occur ]
+ 
+ The patches restructure the way coroutines are allocated and deallocated. 
Coroutines are used extensively in the QMP/HMP protocol implementation, block 
drivers and live migrations; this change is somewhat fundamental to QEMU's 
coroutine infrastructure. If the change were wrong/broken, it could cause:
+ - failing block device drivers, casuing failed boots or broken I/O operations 
within affected guests
+ - incorrect/strange behavior when interacting with a QEMU process over QMP/HMP
+ - failed/broken live migrations
+ - performance regressions as the patch is almost entirely changes in the 
coroutine pool hot path (coroutine_pool_get)
+ 
+ In addition to the above hypotheticals, this change will have the
+ concrete impact of potentially reducing performance for workloads which
+ require many coroutines (as the patches set a hard limit on the size of
+ the coroutine pool, so corountines needed above the pool size are
+ allocated/deallocated on demand). The coroutine pool size is a function
+ of `vm.max_map_count`, so increasing the max map count is the
+ reccommended workaround:
+ 
+ ```
+ sysctl -w vm.max_map_count=1048576
+ ```
+ 
+ It's likely that workloads running into this performance regression
+ would also have been affected by the original bug.
+ 
+ Upstream has two follow-ups to to the refactor,
+ 9352f80cd926fe2dde7c89b93ee33bb0356ff40e and
+ 25bc7d16fa96b0ff881c83ed225ea380fe427c78. 9352f80cd will be included in
+ the upload; 25bc7d16fa fixes a false-positive compiler warning.
  
  [ Other Info ]
  
  QEMU's coroutines can be considered akin to a userspace threading
  implementation. They are used frequently in QEMU's virtual disk drivers.
  
  From upstream, this bug occurs because "per-thread pools can grow to
  tens of thousands of coroutines. Each coroutine causes 2 virtual memory
  areas to be created. Eventually vm.max_map_count is reached and memory-
  related syscalls fail. The per-thread pool sizes are non-uniform and
  depend on past coroutine usage in each thread, so it's possible for one
  thread to have a large pool while another thread's pool is empty."
  
  The new approach "does not leave large numbers of coroutines pooled in a
  thread that may not use them again. In order to perform well it
  amortizes the cost of global pool accesses by working in batches of
  coroutines instead of individual coroutines."
  
  Upstream commit (coroutine: cap per-thread local pool size): 
https://gitlab.com/qemu-project/qemu/-/commit/86a637e4
  Upstream commit (coroutine: reserve 5,000 mappings): 
https://gitlab.com/qemu-project/qemu/-/commit/9352f80c

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2168803

Title:
  QEMU 8.2.2 fails to start guests by exhausting vm.max_map_count with
  coroutines

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/qemu/+bug/2168803/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to