On Mon, Sep 21, 2026 at 10:44:38AM -0700, Sean Christopherson wrote:
> Serialize vCPU creation by holding kvm->lock for the entirety of
> kvm_vm_ioctl_create_vcpu(), and then revert the now-redundant tracking adding
> by commit 97d65b544f48 ("KVM: Check for duplicate vcpu_id as early as
> possible"). I botched the math when justifying the vcpu_ids tracking; it's
> not
> an extra 256 bytes, it's an extra 2048 bytes. Roughly doubling the size of
> "struct kvm" tripped x86's KVM_SANITY_CHECK_VM_STRUCT_SIZE, and obviously
> isn't
> something we want to do in general.
>
> The TL;DR of why it's a-ok to serialize vCPU creation is that no VMM actually
> does parallel vCPU creation. As with so many things, KVM's current behavior
> is
> the result of decades-old cruft, not intentional, deliberate design.
>
> Patches 1-3 are a tangentially related cleanups and bug fixes; I included them
> here because holding kvm->lock for all of vCPU creation allows WARNing if KVM
> attempts to lock all vCPUs if vCPU creation is in-progress (the caller must
> hold kvm->lock).
>
> v2:
> - Tweak patch 1's changelog to clarify that that only x86's manual checks are
> dropped. [Sashiko]
> - Add patches to convert additional arm64 and RISC-V usage to
> kvm_is_vcpu_creation_in_progress(). [Sashiko]
> - Remove acquisition of kvm->lock from s390 and PPC vCPU creation flows.
> [Christian, Sashiko]
> - Add Jean-Christophe's Tested-by to the revert.
>
> v1: https://lore.kernel.org/all/[email protected]
>
> Sean Christopherson (7):
> KVM: Reject attempts to lock all vCPUs if vCPU creation is in-progress
> KVM: arm64: vgic: Rely on vCPU creation check in "trylock all vCPUs"
> KVM: RISC-V: Use kvm_is_vcpu_creation_in_progress() instead of
> open-coded equivalent
> KVM: Protect all of kvm_vm_ioctl_create_vcpu() with kvm->lock
> KVM: Move check for existing vCPU ID to the top of vCPU creation
> Revert "KVM: Check for duplicate vcpu_id as early as possible"
> KVM: WARN if vCPU creation is in-progress when locking all vCPUs
This addresses the Secure TSC splat I reported previously:
https://lore.kernel.org/all/apgj5l7DZGsScsNc@blrnaveerao1/
I booted a SNP guest with Secure TSC enabled and didn't see any lockdep
reports. For what that's worth:
Tested-by: Naveen N Rao (AMD) <[email protected]>
- Naveen