On Mon, Oct 05, 2026 at 01:30:59PM +0530, Amit Machhiwal wrote:
> Hi Gautam,
> 
> Thanks for the patch and working on this.  I have a couple of questions 
> though.
> 
> On 2026/09/21 04:40 PM, Gautam Menghani wrote:
> > A huge number of spurious interrupts can be seen immediately after a KVM
> > on PowerNV guest boots up in XIVE mode.
> > 
> > $ cat /proc/interrupts  | grep SPU
> > SPU:     223705     192439     273526     147623   Spurious interrupts
> > 
> > This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> > Handle pending exceptions on guest entry with MSR_EE"). The root cause
> > is that once LPCR_MER bit is set, it is supposed to be reset by
> > software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> > not exit back to the host until the decrementer expires or there is an
> > hcall, etc. This is because KVM on PowerNV guests have support for
> > native XIVE, so they are not dependent on host for interrupt emulation.
> > Due to this behaviour, a huge number of spurious interrupts are seen
> > since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> > vCPU exits to the host.
> > 
> > Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> > available (currently in case of KVM on PowerNV only), as the XIVE hardware
> > can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> > functionality is not required. This reduces the number of spurious
> > interrupts drastically.
> > 
> > Fixes: ecd10702baae5 ("KVM: PPC: Book3S HV: Handle pending exceptions on 
> > guest entry with MSR_EE")
> > Cc: [email protected] # 6.8+
> > Reported-by: Timothy Pearson <[email protected]>
> > Closes: 
> > https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.javamail.zim...@raptorengineeringinc.com
> > Signed-off-by: Gautam Menghani <[email protected]>
> > ---
> > v3:
> > 1. Continue the use of LPCR_MER when kernel-irqchip=off (Sashiko)
> > 
> > v2:
> > 1. Handle the case where xive_interrupt_pending() is true and also the
> > external exception bit is set. (Narayana)
> > 
> >  arch/powerpc/include/asm/kvm_ppc.h | 7 +++++++
> >  arch/powerpc/kvm/book3s_hv.c       | 2 +-
> >  2 files changed, 8 insertions(+), 1 deletion(-)
> > 
> > diff --git a/arch/powerpc/include/asm/kvm_ppc.h 
> > b/arch/powerpc/include/asm/kvm_ppc.h
> > index 169ea6a7fbad..580ad2548c2b 100644
> > --- a/arch/powerpc/include/asm/kvm_ppc.h
> > +++ b/arch/powerpc/include/asm/kvm_ppc.h
> > @@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu 
> > *vcpu)
> >     return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
> >  }
> >  
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
> > +{
> > +   return kvm->arch.xive_devices.native;
> > +}
> > +
> >  extern int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> >                                        struct kvm_vcpu *vcpu, u32 cpu);
> >  extern void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu);
> > @@ -782,6 +787,8 @@ static inline bool kvmppc_xive_rearm_escalation(struct 
> > kvm_vcpu *vcpu) { return
> >  
> >  static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> >     { return 0; }
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm) { return 
> > false; }
> > +
> >  static inline int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> >                       struct kvm_vcpu *vcpu, u32 cpu) { return -EBUSY; }
> >  static inline void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu) 
> > { }
> > diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
> > index dbac3573b2c8..cb2bb29451a7 100644
> > --- a/arch/powerpc/kvm/book3s_hv.c
> > +++ b/arch/powerpc/kvm/book3s_hv.c
> > @@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 
> > time_limit,
> >                     if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & 
> > MSR_EE))
> >                             kvmppc_inject_interrupt_hv(vcpu,
> >                                                        
> > BOOK3S_INTERRUPT_EXTERNAL, 0);
> > -                   else
> > +                   else if (!kvmppc_xive_native_enabled(vcpu->kvm))
> >                             lpcr |= LPCR_MER;
> 
> 1. What happens when an L1 KVM guest is booted with (native) XIVE and then the
>    guest is rebooted with `xive=off` i.e., with XICS?  Will we stop setting
>    LPCR_MER as kvm->arch.xive_devices.native would still be set?

Yes, that's a good catch. kvm->arch.xive_devices.native is cleared only
when vm is destroyed. So this case has to be handled.

> 2. What happends when an L1 KVM guest is booted with XIVE and then we kexec 
> into
>    a new kernel with `xive=off`?  You did mention in your other reply that 
> with
>    kexec, `xive=off` is ignored currently but IMO, we would want to understand
>    where this limiation lies and fix that if need be.  I do understand we are
>    trying to fix spurious interrupts problem with XIVE in this patch and this
>    particular problem can be taken separately but it worth investigating.

Yes right, the kexec case is to be understood, but I'll take it up separately.
Meanwhile, I'll send a v4 to fix the reboot case.

> 
> Thanks,
> Amit

Reply via email to