** Description changed:

  [ Impact ]
  
  QEMU 8.2.2's coroutine pool implementation can hit the Linux
  vm.max_map_count limit (`65530` in Jammy, `1048576` in Noble), causing
  QEMU to abort with "failed to allocate memory for stack" or "failed to
  set up stack guard page" during coroutine creation.
  
- Each coroutine allocates two VMAs with mmap(2). When a coroutine
- terminates it is cached in a global "release pool"; once that is full,
- terminated coroutines are cached in the terminating thread's "alloc
- pool" (up to the max pool size, per thread). When the thread-local pool
- is empty, the global pool is moved to the thread-local pool. Pooled
- coroutines are only freed once both pools are full.
- 
- The theoretical maximum of coroutines in each thread-local pool is given
- by: (see hw/block/virtio-blk.c:1638)
+ Each coroutine creates two anonymous mappings with mmap(2). When a
+ coroutine terminates it is cached in a global "release pool"; once that
+ is full (2 x max pool size), terminated coroutines are cached in a
+ thread-local "alloc pool" (up to the max pool size, per thread). When
+ the thread-local pool is empty, the global pool is moved to the thread-
+ local pool. Pooled coroutines are only freed once both pools are full.
+ 
+ In (unpatched) QEMU 8.2.2 the theoretical maximum of idle coroutines in
+ each thread-local pool is given by: (see hw/block/virtio-blk.c:1638)
  
  64 + Σ virtio-blk devices (num_queues_i × queue_size_i / 2)
  
  Defaults:
  
  queue_size_i = 256 (hw/block/virtio-blk.c:1722)
  num_queues_i = vCPUs (hw/virtio/virtio-pci.c:2466)
  
- Assume the virtio-blk defaults for a VM with one disk and 32 vCPUs and
- you get 4160 coroutines _per IOthread_, plus 8320 for the global pool
- (2x the max size for a thread-local pool). Assuming the same number of
- IOThreads as vCPUs (x32) gives a theoretical maximum of 141440 cached
- coroutines. x2 for two VMAs per coroutine for a grand theoretical max of
- 282880 VMAs.
+ The global pool is twice as large as the max pool size, so:
+ 
+ theoretical maximum mappings = 2 × (iothreads + 2) × thread-local max
+ 
+ A VM with 32 vCPUs and two virtio-blk disks gets a theoretical maximum
+ mappings of 49920; add a third disk or a few more vCPUs and it is over
+ 65530 for older kernels.
  
  In practice the number of coroutines that will actually be allocated is
- much lower than this, since in order to reach this maximum all of them
- would need to be active at once. Many small IOs all at once drain both
- pools, causing more coroutines to be allocated.
+ typically much lower than this, since the total number of allocated
+ coroutines is a high water mark of all concurrently active coroutines.
+ Many small IOs all at once drain both pools, causing more coroutines to
+ be allocated.
  
  This behavior was introduced in upstream 4c41c69e05fe28c (v7.0.0-rc0)
  and therefore only affects Ubuntu 24.04.
  
  The user who reported this experienced crashes while running Noble's
  QEMU in a container on a 5.15 kernel (with the old
  `vm.max_map_count=65530`). Even with a relatively modest VM
- configuration, the theoretical max of coroutine VMAs is substantially
- higher than this limit.
- 
- Crashes are expected to be significantly less likely with a Noble kernel
- since reaching the default max_map_count would require the pool's
- theoretical max size to be increased (using 128+ IOThreads or adding
- additional virtio-blk devices). Larger VMs might blow past the higher
- default in Noble's kernel.
- 
- While the issue can be worked around by adjusting the sysctl, this is
- still a resource leak that leaves significant numbers of unused
- coroutines allocated but unusable. Enforcing that the limit of
- coroutines be tied to the `vm.max_map_count` hardens VMs against
- "bursty" IO causing unexpected crashes.
+ configuration (36 vCPUs, 2 disks), the theoretical max of coroutine VMAs
+ is higher than this limit.
+ 
+ Crashes are expected to be significantly less likely with Noble's sysctl
+ config since reaching the default max_map_count would require the pool's
+ theoretical max size to be increased substantially (using 128+ IOThreads
+ or adding many additional virtio-blk devices). Crashes with this
+ configuration have not been reported.
+ 
+ While crashes can be worked around by adjusting the sysctl, this is
+ still a resource leak that can leave a large number of unused coroutines
+ allocated but unusable. 86a637e4810 allows for better reuse of of those
+ cached idle coroutines.
  
  [ Test Plan ]
+ 
+ 1. Run the upstream iotests (in a virtual machine) to exercise coroutines (not
+ all of these tests are run during package build):
+ 
+ ```sh
+ gbp pq switch
+ sudo apt build-dep ./
+ sudo apt install seabios ipxe-qemu
+ 
+ make -f debian/rules configure-qemu
+ 
+ # Binaries needed for iotests
+ ninja -C b/qemu modules
+ make -f debian/rules build-x86-optionrom
+ cp -a b/optionrom/*.bin /usr/share/seabios/* /usr/lib/ipxe/qemu/* 
b/qemu/qemu-bundle/usr/share/qemu/
+ 
+ make -C b/qemu/ check-block
+ cd b/qemu/
+ ./tests/qemu-iotests/check -qcow2
+ # Expected failures 151 281 307 graph-changes-while-io
+ ```
+ 
+ 2. Run the Ubuntu Server team's qemu-migration-test [1]
+ 
+ 3. The following procedure reproduces mmap exhaustion:
  
  ```
  lxc launch ubuntu:noble --vm -c limits.cpu=8 -c limits.memory=16GiB n0
  ```
  
  In the VM:
  ```sh
  sudo apt install libvirt-daemon-system virtinst
  
  wget 
https://cloud-images.ubuntu.com/daily/server/noble/current/noble-server-cloudimg-amd64.img
  sudo cp noble-server-cloudimg-amd64.img 
/var/lib/libvirt/images/testvm-root.qcow2
  sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data.raw 10G
  sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data2.raw 10G
  sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data3.raw 10G
  sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data4.raw 10G
  
  cat > user-data <<EOF
  #cloud-config
  password: ubuntu
  chpasswd:
    expire: false
  ssh_pwauth: true
  package_update: true
  packages:
    - fio
  EOF
  
  sudo virt-install \
    -n testvm \
    --os-variant=ubuntu24.04 \
    --ram=8192 --vcpus=8 \
    --iothreads 8 \
    --import \
    --disk 
path=/var/lib/libvirt/images/testvm-root.qcow2,bus=virtio,cache=writethrough \
    --disk 
path=/var/lib/libvirt/images/testvm-data.raw,bus=virtio,cache=none,driver.io=io_uring
 \
    --disk 
path=/var/lib/libvirt/images/testvm-data2.raw,bus=virtio,cache=none,driver.io=io_uring
 \
    --disk 
path=/var/lib/libvirt/images/testvm-data3.raw,bus=virtio,cache=none,driver.io=io_uring
 \
    --disk 
path=/var/lib/libvirt/images/testvm-data4.raw,bus=virtio,cache=none,driver.io=io_uring
 \
    --graphics none \
    --network network=default \
    --cloud-init user-data=user-data
  
  # Ctrl+Shift+a+]
  
  sudo virsh shutdown testvm
  ```
  
  Limit write throughput in the block layer so that writes allocate a coroutine
  but aren't able to complete right away:
  ```sh
  for disk in vdb vdc vdd vde; do
    sudo virsh blkdeviotune testvm "${disk}" --total-iops-sec 2000 --current
  done
  ```
  
  Use `sudo virsh edit testvm` to pin each disk to its own iothread; otherwise
  all IO will occur on the main thread. Add `iothread='1'` to the `<driver>`
  section of each 'testvm-data' disk, incrementing the number for each disk.
  
  ```sh
  # Reduce the max map count by a factor of four to match the reduction in
  # used IOThreads in the test environment
  sudo sysctl -w vm.max_map_count=16384
  
  sudo virsh start testvm
  sudo virsh console testvm
  ```
  
  On the host, you can monitor the maps with:
  ```sh
  pid=$(pidof qemu-system-x86_64)
  while sleep 1; do sudo grep -c '' "/proc/${pid}/maps"; done
  ```
  
  In the testvm:
  ```sh
  while true; do
    for dev in vdb vdc vdd vde; do
      sudo fio "--name=/dev/${dev}" \
        --ioengine=io_uring \
        --direct=1 \
        --runtime=10 \
        --time_based=1 \
        --rw=randread \
        --bs=4k \
        --iodepth=128 \
        --numjobs=32 \
        --group_reporting \
        "--filename=/dev/${dev}"
    done
  done
  ```
  
  Patched behavior:
  
  /proc/pid/maps levels off and stays consistent through the fio run. (With 
these
  parameters we observed approximately 5376).
  
  Unpatched behavior:
  
  /proc/pid/maps increases with each fio run until the VM crashes with one
  of:
  
  - failed to allocate memory for stack: Cannot allocate memory
  - failed to set up stack guard page: Cannot allocate memory
  
- Performance regression test (results should be provided both before & after):
+ 4. Performance regression test (results should be provided both before &
+ after):
+ 
  ```sh
  for bs in 4k 128k; do
    for iodepth in 1 8 16 32 64; do
-     sudo fio --output-format=json --output=fio-$bs-$iodepth.json \
-       --ioengine=libaio --direct=1 --numjobs=$(nproc) \
-       --runtime=30 --ramp_time=5 --time_based=1 \
-       --rw=randread --bs=$bs --iodepth=$iodepth \
-       --name=job1 --filename=/dev/vdb
+     sudo fio --output-format=json \
+       --output=fio-$bs-$iodepth.json \
+       --ioengine=libaio \
+       --direct=1 \
+       --numjobs=$(nproc) \
+       --runtime=30 \
+       --ramp_time=5 \
+       --time_based=1 \
+       --rw=randread \
+       --bs=$bs \
+       --iodepth=$iodepth \
+       --name=job1 \
+       --filename=/dev/vdb
    done
  done
  ```
+ 
+ [1] https://code.launchpad.net/~ubuntu-server/ubuntu/+source/qemu-
+ migration-test/+git/qemu-migration-test
  
  [ Where problems could occur ]
  
  The patches restructure the way coroutines are allocated and deallocated. 
Coroutines are used extensively in the QMP/HMP protocol implementation, block 
drivers and live migrations; this change is somewhat fundamental to QEMU's 
coroutine infrastructure. If the change were wrong/broken, it could cause:
  - failing block device drivers, causing failed boots or broken I/O operations 
within affected guests
  - incorrect/strange behavior when interacting with a QEMU process over QMP/HMP
  - failed/broken live migrations
- - performance regressions as the patch is almost entirely changes in the 
coroutine pool hot path (coroutine_pool_get)
- 
- In addition to the above hypotheticals, this change will have the
- concrete impact of potentially reducing performance for workloads which
- require many coroutines (as the patches set a hard limit on the size of
- the coroutine pool, so coroutines needed above the pool size are
- allocated/deallocated on demand). The coroutine pool size is a function
- of `vm.max_map_count`, so increasing the max map count is the
- recommended workaround:
- 
- ```
- sysctl -w vm.max_map_count=1048576
- ```
- 
- It's likely that workloads running into this performance regression
- would also have been affected by the original bug.
+ - performance regressions as the patch introduces a lock in the coroutine 
pool hot path (coroutine_pool_get)
  
  Upstream has two follow-ups to the refactor,
  9352f80cd926fe2dde7c89b93ee33bb0356ff40e and
  25bc7d16fa96b0ff881c83ed225ea380fe427c78. 9352f80cd will be included in
  the upload; 25bc7d16fa fixes a false-positive compiler warning.
  
+ 86a637e4 parses the contents of `/proc/sys/vm/max_map_count` incorrectly
+ [1]. This means that without Hector's linked follow-up, the patch
+ removes any limitation on the maximum size of the coroutine idle pool.
+ 
+ The introduction of a read() from procfs will introduce an apparmor
+ denial for VMs using libvirt's AppArmor security driver [2]. If AppArmor
+ is active for a VM, this change will remove any limitation of the
+ maximum size of the coroutine idle pool.
+ 
+ If [1] and [2] are also applied, workloads which require many coroutines
+ may experience a performance regression (as the patches set a hard limit
+ on the size of the coroutine pool, so coroutines needed above the pool
+ size are allocated/deallocated on demand). The coroutine pool size is a
+ function of `vm.max_map_count`, so increasing the max map count is the
+ recommended workaround:
+ 
+ ```
+ sysctl -w vm.max_map_count=1048576
+ ```
+ 
+ It's likely that workloads running into this performance regression
+ would also have been affected by the original bug.
+ 
+ [1] https://lists.gnu.org/archive/html/qemu-devel/2026-10/msg01192.html
+ [2] 
https://gitlab.com/libvirt/libvirt/-/commit/85e07fb1ceee7943879f8a374cabfa8ab858a3c6
+ 
  [ Other Info ]
  
- QEMU's coroutines can be considered akin to a userspace threading
- implementation. They are used frequently in QEMU's virtual disk drivers.
+ QEMU's coroutines are a userspace threading implementation. They are
+ used frequently in QEMU's virtual disk drivers.
  
  Noble's QEMU 8.2.2 keeps a thread-local list of coroutines `alloc_pool`
  (util/qemu-coroutine.c:40); this list can grow to thousands or even tens
  of thousands of coroutines _for each thread_. Each coroutine has two
  VMAs created with mmap(2). If there are enough threads available to the
  QEMU process, the number of coroutines may cause the total number of
  mappings to exceed `vm.max_map_count`, causing the QEMU process to fail
  to allocate memory and crash.
  
  From upstream, this bug occurs because "per-thread pools can grow to
  tens of thousands of coroutines. Each coroutine causes 2 virtual memory
  areas to be created. Eventually vm.max_map_count is reached and memory-
  related syscalls fail. The per-thread pool sizes are non-uniform and
  depend on past coroutine usage in each thread, so it's possible for one
  thread to have a large pool while another thread's pool is empty."
  
  The new approach "does not leave large numbers of coroutines pooled in a
  thread that may not use them again. In order to perform well it
  amortizes the cost of global pool accesses by working in batches of
  coroutines instead of individual coroutines."
  
  Upstream commit (coroutine: cap per-thread local pool size): 
https://gitlab.com/qemu-project/qemu/-/commit/86a637e4
  Upstream commit (coroutine: reserve 5,000 mappings): 
https://gitlab.com/qemu-project/qemu/-/commit/9352f80c
  Upstream commit (util/coroutine: fix -Werror=maybe-uninitialized 
false-positive): https://gitlab.com/qemu-project/qemu/-/commit/25bc7d
+ 
+ Additionally, the following bugs were identified in 86a637e4 which
+ require follow-ups for the patch to work as intended:
+ 
+ Missing AppArmor rule: 
https://gitlab.com/libvirt/libvirt/-/commit/85e07fb1ceee7943879f8a374cabfa8ab858a3c6
+ Incorrect string parsing: 
https://lists.gnu.org/archive/html/qemu-devel/2026-10/msg01192.html

** Description changed:

  [ Impact ]
  
  QEMU 8.2.2's coroutine pool implementation can hit the Linux
  vm.max_map_count limit (`65530` in Jammy, `1048576` in Noble), causing
  QEMU to abort with "failed to allocate memory for stack" or "failed to
  set up stack guard page" during coroutine creation.
  
  Each coroutine creates two anonymous mappings with mmap(2). When a
  coroutine terminates it is cached in a global "release pool"; once that
  is full (2 x max pool size), terminated coroutines are cached in a
  thread-local "alloc pool" (up to the max pool size, per thread). When
  the thread-local pool is empty, the global pool is moved to the thread-
  local pool. Pooled coroutines are only freed once both pools are full.
  
  In (unpatched) QEMU 8.2.2 the theoretical maximum of idle coroutines in
  each thread-local pool is given by: (see hw/block/virtio-blk.c:1638)
  
  64 + Σ virtio-blk devices (num_queues_i × queue_size_i / 2)
  
  Defaults:
  
  queue_size_i = 256 (hw/block/virtio-blk.c:1722)
  num_queues_i = vCPUs (hw/virtio/virtio-pci.c:2466)
  
  The global pool is twice as large as the max pool size, so:
  
  theoretical maximum mappings = 2 × (iothreads + 2) × thread-local max
  
  A VM with 32 vCPUs and two virtio-blk disks gets a theoretical maximum
  mappings of 49920; add a third disk or a few more vCPUs and it is over
  65530 for older kernels.
  
  In practice the number of coroutines that will actually be allocated is
  typically much lower than this, since the total number of allocated
  coroutines is a high water mark of all concurrently active coroutines.
  Many small IOs all at once drain both pools, causing more coroutines to
  be allocated.
  
  This behavior was introduced in upstream 4c41c69e05fe28c (v7.0.0-rc0)
  and therefore only affects Ubuntu 24.04.
  
  The user who reported this experienced crashes while running Noble's
  QEMU in a container on a 5.15 kernel (with the old
  `vm.max_map_count=65530`). Even with a relatively modest VM
- configuration (36 vCPUs, 2 disks), the theoretical max of coroutine VMAs
+ configuration (36 vCPUs, 3 disks), the theoretical max of coroutine VMAs
  is higher than this limit.
  
  Crashes are expected to be significantly less likely with Noble's sysctl
  config since reaching the default max_map_count would require the pool's
  theoretical max size to be increased substantially (using 128+ IOThreads
  or adding many additional virtio-blk devices). Crashes with this
  configuration have not been reported.
  
  While crashes can be worked around by adjusting the sysctl, this is
  still a resource leak that can leave a large number of unused coroutines
  allocated but unusable. 86a637e4810 allows for better reuse of of those
  cached idle coroutines.
  
  [ Test Plan ]
  
  1. Run the upstream iotests (in a virtual machine) to exercise coroutines (not
  all of these tests are run during package build):
  
  ```sh
  gbp pq switch
  sudo apt build-dep ./
  sudo apt install seabios ipxe-qemu
  
  make -f debian/rules configure-qemu
  
  # Binaries needed for iotests
  ninja -C b/qemu modules
  make -f debian/rules build-x86-optionrom
  cp -a b/optionrom/*.bin /usr/share/seabios/* /usr/lib/ipxe/qemu/* 
b/qemu/qemu-bundle/usr/share/qemu/
  
  make -C b/qemu/ check-block
  cd b/qemu/
  ./tests/qemu-iotests/check -qcow2
  # Expected failures 151 281 307 graph-changes-while-io
  ```
  
  2. Run the Ubuntu Server team's qemu-migration-test [1]
  
  3. The following procedure reproduces mmap exhaustion:
  
  ```
  lxc launch ubuntu:noble --vm -c limits.cpu=8 -c limits.memory=16GiB n0
  ```
  
  In the VM:
  ```sh
  sudo apt install libvirt-daemon-system virtinst
  
  wget 
https://cloud-images.ubuntu.com/daily/server/noble/current/noble-server-cloudimg-amd64.img
  sudo cp noble-server-cloudimg-amd64.img 
/var/lib/libvirt/images/testvm-root.qcow2
  sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data.raw 10G
  sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data2.raw 10G
  sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data3.raw 10G
  sudo qemu-img create -f raw /var/lib/libvirt/images/testvm-data4.raw 10G
  
  cat > user-data <<EOF
  #cloud-config
  password: ubuntu
  chpasswd:
-   expire: false
+   expire: false
  ssh_pwauth: true
  package_update: true
  packages:
-   - fio
+   - fio
  EOF
  
  sudo virt-install \
-   -n testvm \
-   --os-variant=ubuntu24.04 \
-   --ram=8192 --vcpus=8 \
-   --iothreads 8 \
-   --import \
-   --disk 
path=/var/lib/libvirt/images/testvm-root.qcow2,bus=virtio,cache=writethrough \
-   --disk 
path=/var/lib/libvirt/images/testvm-data.raw,bus=virtio,cache=none,driver.io=io_uring
 \
-   --disk 
path=/var/lib/libvirt/images/testvm-data2.raw,bus=virtio,cache=none,driver.io=io_uring
 \
-   --disk 
path=/var/lib/libvirt/images/testvm-data3.raw,bus=virtio,cache=none,driver.io=io_uring
 \
-   --disk 
path=/var/lib/libvirt/images/testvm-data4.raw,bus=virtio,cache=none,driver.io=io_uring
 \
-   --graphics none \
-   --network network=default \
-   --cloud-init user-data=user-data
+   -n testvm \
+   --os-variant=ubuntu24.04 \
+   --ram=8192 --vcpus=8 \
+   --iothreads 8 \
+   --import \
+   --disk 
path=/var/lib/libvirt/images/testvm-root.qcow2,bus=virtio,cache=writethrough \
+   --disk 
path=/var/lib/libvirt/images/testvm-data.raw,bus=virtio,cache=none,driver.io=io_uring
 \
+   --disk 
path=/var/lib/libvirt/images/testvm-data2.raw,bus=virtio,cache=none,driver.io=io_uring
 \
+   --disk 
path=/var/lib/libvirt/images/testvm-data3.raw,bus=virtio,cache=none,driver.io=io_uring
 \
+   --disk 
path=/var/lib/libvirt/images/testvm-data4.raw,bus=virtio,cache=none,driver.io=io_uring
 \
+   --graphics none \
+   --network network=default \
+   --cloud-init user-data=user-data
  
  # Ctrl+Shift+a+]
  
  sudo virsh shutdown testvm
  ```
  
  Limit write throughput in the block layer so that writes allocate a coroutine
  but aren't able to complete right away:
  ```sh
  for disk in vdb vdc vdd vde; do
-   sudo virsh blkdeviotune testvm "${disk}" --total-iops-sec 2000 --current
+   sudo virsh blkdeviotune testvm "${disk}" --total-iops-sec 2000 --current
  done
  ```
  
  Use `sudo virsh edit testvm` to pin each disk to its own iothread; otherwise
  all IO will occur on the main thread. Add `iothread='1'` to the `<driver>`
  section of each 'testvm-data' disk, incrementing the number for each disk.
  
  ```sh
  # Reduce the max map count by a factor of four to match the reduction in
  # used IOThreads in the test environment
  sudo sysctl -w vm.max_map_count=16384
  
  sudo virsh start testvm
  sudo virsh console testvm
  ```
  
  On the host, you can monitor the maps with:
  ```sh
  pid=$(pidof qemu-system-x86_64)
  while sleep 1; do sudo grep -c '' "/proc/${pid}/maps"; done
  ```
  
  In the testvm:
  ```sh
  while true; do
-   for dev in vdb vdc vdd vde; do
-     sudo fio "--name=/dev/${dev}" \
-       --ioengine=io_uring \
-       --direct=1 \
-       --runtime=10 \
-       --time_based=1 \
-       --rw=randread \
-       --bs=4k \
-       --iodepth=128 \
-       --numjobs=32 \
-       --group_reporting \
-       "--filename=/dev/${dev}"
-   done
+   for dev in vdb vdc vdd vde; do
+     sudo fio "--name=/dev/${dev}" \
+       --ioengine=io_uring \
+       --direct=1 \
+       --runtime=10 \
+       --time_based=1 \
+       --rw=randread \
+       --bs=4k \
+       --iodepth=128 \
+       --numjobs=32 \
+       --group_reporting \
+       "--filename=/dev/${dev}"
+   done
  done
  ```
  
  Patched behavior:
  
  /proc/pid/maps levels off and stays consistent through the fio run. (With 
these
  parameters we observed approximately 5376).
  
  Unpatched behavior:
  
  /proc/pid/maps increases with each fio run until the VM crashes with one
  of:
  
  - failed to allocate memory for stack: Cannot allocate memory
  - failed to set up stack guard page: Cannot allocate memory
  
  4. Performance regression test (results should be provided both before &
  after):
  
  ```sh
  for bs in 4k 128k; do
-   for iodepth in 1 8 16 32 64; do
-     sudo fio --output-format=json \
-       --output=fio-$bs-$iodepth.json \
-       --ioengine=libaio \
-       --direct=1 \
-       --numjobs=$(nproc) \
-       --runtime=30 \
-       --ramp_time=5 \
-       --time_based=1 \
-       --rw=randread \
-       --bs=$bs \
-       --iodepth=$iodepth \
-       --name=job1 \
-       --filename=/dev/vdb
-   done
+   for iodepth in 1 8 16 32 64; do
+     sudo fio --output-format=json \
+       --output=fio-$bs-$iodepth.json \
+       --ioengine=libaio \
+       --direct=1 \
+       --numjobs=$(nproc) \
+       --runtime=30 \
+       --ramp_time=5 \
+       --time_based=1 \
+       --rw=randread \
+       --bs=$bs \
+       --iodepth=$iodepth \
+       --name=job1 \
+       --filename=/dev/vdb
+   done
  done
  ```
  
  [1] https://code.launchpad.net/~ubuntu-server/ubuntu/+source/qemu-
  migration-test/+git/qemu-migration-test
  
  [ Where problems could occur ]
  
  The patches restructure the way coroutines are allocated and deallocated. 
Coroutines are used extensively in the QMP/HMP protocol implementation, block 
drivers and live migrations; this change is somewhat fundamental to QEMU's 
coroutine infrastructure. If the change were wrong/broken, it could cause:
  - failing block device drivers, causing failed boots or broken I/O operations 
within affected guests
  - incorrect/strange behavior when interacting with a QEMU process over QMP/HMP
  - failed/broken live migrations
  - performance regressions as the patch introduces a lock in the coroutine 
pool hot path (coroutine_pool_get)
  
  Upstream has two follow-ups to the refactor,
  9352f80cd926fe2dde7c89b93ee33bb0356ff40e and
  25bc7d16fa96b0ff881c83ed225ea380fe427c78. 9352f80cd will be included in
  the upload; 25bc7d16fa fixes a false-positive compiler warning.
  
  86a637e4 parses the contents of `/proc/sys/vm/max_map_count` incorrectly
  [1]. This means that without Hector's linked follow-up, the patch
  removes any limitation on the maximum size of the coroutine idle pool.
  
  The introduction of a read() from procfs will introduce an apparmor
  denial for VMs using libvirt's AppArmor security driver [2]. If AppArmor
  is active for a VM, this change will remove any limitation of the
  maximum size of the coroutine idle pool.
  
  If [1] and [2] are also applied, workloads which require many coroutines
  may experience a performance regression (as the patches set a hard limit
  on the size of the coroutine pool, so coroutines needed above the pool
  size are allocated/deallocated on demand). The coroutine pool size is a
  function of `vm.max_map_count`, so increasing the max map count is the
  recommended workaround:
  
  ```
  sysctl -w vm.max_map_count=1048576
  ```
  
  It's likely that workloads running into this performance regression
  would also have been affected by the original bug.
  
  [1] https://lists.gnu.org/archive/html/qemu-devel/2026-10/msg01192.html
  [2] 
https://gitlab.com/libvirt/libvirt/-/commit/85e07fb1ceee7943879f8a374cabfa8ab858a3c6
  
  [ Other Info ]
  
  QEMU's coroutines are a userspace threading implementation. They are
  used frequently in QEMU's virtual disk drivers.
  
  Noble's QEMU 8.2.2 keeps a thread-local list of coroutines `alloc_pool`
  (util/qemu-coroutine.c:40); this list can grow to thousands or even tens
  of thousands of coroutines _for each thread_. Each coroutine has two
  VMAs created with mmap(2). If there are enough threads available to the
  QEMU process, the number of coroutines may cause the total number of
  mappings to exceed `vm.max_map_count`, causing the QEMU process to fail
  to allocate memory and crash.
  
  From upstream, this bug occurs because "per-thread pools can grow to
  tens of thousands of coroutines. Each coroutine causes 2 virtual memory
  areas to be created. Eventually vm.max_map_count is reached and memory-
  related syscalls fail. The per-thread pool sizes are non-uniform and
  depend on past coroutine usage in each thread, so it's possible for one
  thread to have a large pool while another thread's pool is empty."
  
  The new approach "does not leave large numbers of coroutines pooled in a
  thread that may not use them again. In order to perform well it
  amortizes the cost of global pool accesses by working in batches of
  coroutines instead of individual coroutines."
  
  Upstream commit (coroutine: cap per-thread local pool size): 
https://gitlab.com/qemu-project/qemu/-/commit/86a637e4
  Upstream commit (coroutine: reserve 5,000 mappings): 
https://gitlab.com/qemu-project/qemu/-/commit/9352f80c
  Upstream commit (util/coroutine: fix -Werror=maybe-uninitialized 
false-positive): https://gitlab.com/qemu-project/qemu/-/commit/25bc7d
  
  Additionally, the following bugs were identified in 86a637e4 which
  require follow-ups for the patch to work as intended:
  
  Missing AppArmor rule: 
https://gitlab.com/libvirt/libvirt/-/commit/85e07fb1ceee7943879f8a374cabfa8ab858a3c6
  Incorrect string parsing: 
https://lists.gnu.org/archive/html/qemu-devel/2026-10/msg01192.html

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2168803

Title:
  QEMU 8.2.2 guests crash by exhausting vm.max_map_count with coroutines

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/qemu/+bug/2168803/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to