On Mon, Sep 14, 2026, Masami Hiramatsu (Google) wrote:
> From: Masami Hiramatsu (Google) <[email protected]>
> 
> When KVM enters a guest OS, host hardware breakpoints are disabled
> before running the guest. However, an NMI can occur during guest
> execution, where local_db_save() or arch_install_hw_breakpoint()
> can be invoked.
> 
> In particular, if local_db_save() or arch_install_hw_breakpoint()
> is executed from NMI, hardware DR7 can be modified or restored with
> host breakpoint settings, leaking host breakpoints into the guest OS
> or clobbering the guest's debug registers.

Not for local_db_save(), at least not AFAICT.  On VM-Exit, both Intel and AMD
purge DR7, i.e. load 0x400, so local_db_save() => local_db_restore() is more or
less a nop.  Even if that weren't the case, actually saving/restoring DR7 would
be the right thing to do, in any context.

> Introduce a per-CPU flag, cpu_dr_in_guest, to indicate that the CPU
> is executing in guest mode. Set this flag in vcpu_enter_guest()
> during entering the guest with disabling host breakpoints.
> If this flag is set, local_db_save() and local_db_restore() return
> immediately, and arch_install_hw_breakpoint() returns an error.
> In addition, protect cpu_dr_in_guest in within_cpu_entry() to
> prevent recursive #DB exceptions.
> 
> Fixes: f85d40160691 ("KVM: X86: Disable hardware breakpoints unconditionally 
> before kvm_x86->run()")
> Assisted-by: Antigravity:gemini-3.8-flash
> Signed-off-by: Masami Hiramatsu (Google) <[email protected]>
> ---
> Changes in v16:
>  - Newly added.
> ---
>  arch/x86/include/asm/debugreg.h |    6 ++++++
>  arch/x86/kernel/hw_breakpoint.c |   10 ++++++++++
>  arch/x86/kvm/x86.c              |    7 +++++++
>  3 files changed, 23 insertions(+)
> 
> diff --git a/arch/x86/include/asm/debugreg.h b/arch/x86/include/asm/debugreg.h
> index 854d82b88ff4..50f830972698 100644
> --- a/arch/x86/include/asm/debugreg.h
> +++ b/arch/x86/include/asm/debugreg.h
> @@ -18,6 +18,7 @@
>  #define DR7_FIXED_1  0x00000400
>  
>  DECLARE_PER_CPU(unsigned long, cpu_dr7);
> +DECLARE_PER_CPU(bool, cpu_dr_in_guest);
>  
>  #ifndef CONFIG_PARAVIRT_XXL
>  /*
> @@ -129,6 +130,9 @@ static __always_inline unsigned long local_db_save(void)
>  {
>       unsigned long dr7;
>  
> +     if (this_cpu_read(cpu_dr_in_guest))
> +             return 0;

This is broken.  If an NMI hits between KVM writing cpu_dr_in_guest and clearing
DR7, and there are active breakpoints, then local_db_save() won't disable 
breakpoints
as it should, and the relevant code in exc_nmi() will run with breakpoints 
enabled.

        kvm_load_xfeatures(vcpu, true);

        this_cpu_write(cpu_dr_in_guest, true);
        barrier();

  <NMI here is problematic>

        if (unlikely(vcpu->arch.switch_db_regs &&
                     !(vcpu->arch.switch_db_regs & KVM_DEBUGREG_AUTO_SWITCH))) {
                set_debugreg(DR7_FIXED_1, 7);
                set_debugreg(vcpu->arch.eff_db[0], 0);

> +
>       if (cpu_feature_enabled(X86_FEATURE_HYPERVISOR) && 
> !hw_breakpoint_active())
>               return 0;
>  
> @@ -157,6 +161,8 @@ static __always_inline void local_db_restore(unsigned 
> long dr7)
>        * not be good.
>        */
>       barrier();
> +     if (this_cpu_read(cpu_dr_in_guest))
> +             return;
>       if (dr7)
>               set_debugreg(dr7, 7);
>  }
> diff --git a/arch/x86/kernel/hw_breakpoint.c b/arch/x86/kernel/hw_breakpoint.c
> index f846c15f21ca..68de7ed79d88 100644
> --- a/arch/x86/kernel/hw_breakpoint.c
> +++ b/arch/x86/kernel/hw_breakpoint.c
> @@ -40,6 +40,9 @@
>  DEFINE_PER_CPU(unsigned long, cpu_dr7);
>  EXPORT_PER_CPU_SYMBOL(cpu_dr7);
>  
> +DEFINE_PER_CPU(bool, cpu_dr_in_guest);
> +EXPORT_PER_CPU_SYMBOL_GPL(cpu_dr_in_guest);
> +
>  /* Per cpu debug address registers values */
>  static DEFINE_PER_CPU(unsigned long, cpu_debugreg[HBP_NUM]);
>  
> @@ -102,6 +105,9 @@ int arch_install_hw_breakpoint(struct perf_event *bp)
>  
>       lockdep_assert_irqs_disabled();
>  
> +     if (this_cpu_read(cpu_dr_in_guest))
> +             return -EBUSY;

If we decide this is how to fix arch_install_hw_breakpoint() clobbering DRs from
NMI context, I would rather have more generic flag to tell perf that KVM is 
about
to enter the guest, e.g. so that we don't have to separately solve the same 
problem
for other perf events:

https://lore.kernel.org/all/[email protected]

Reply via email to