Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit

From: Gautam Menghani

Date: Mon Oct 05 2026 - 04:31:56 EST


On Mon, Oct 05, 2026 at 01:30:59PM +0530, Amit Machhiwal wrote:
> Hi Gautam,
>
> Thanks for the patch and working on this. I have a couple of questions though.
>
> On 2026/09/21 04:40 PM, Gautam Menghani wrote:
> > A huge number of spurious interrupts can be seen immediately after a KVM
> > on PowerNV guest boots up in XIVE mode.
> >
> > $ cat /proc/interrupts | grep SPU
> > SPU: 223705 192439 273526 147623 Spurious interrupts
> >
> > This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> > Handle pending exceptions on guest entry with MSR_EE"). The root cause
> > is that once LPCR_MER bit is set, it is supposed to be reset by
> > software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> > not exit back to the host until the decrementer expires or there is an
> > hcall, etc. This is because KVM on PowerNV guests have support for
> > native XIVE, so they are not dependent on host for interrupt emulation.
> > Due to this behaviour, a huge number of spurious interrupts are seen
> > since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> > vCPU exits to the host.
> >
> > Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> > available (currently in case of KVM on PowerNV only), as the XIVE hardware
> > can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> > functionality is not required. This reduces the number of spurious
> > interrupts drastically.
> >
> > Fixes: ecd10702baae5 ("KVM: PPC: Book3S HV: Handle pending exceptions on guest entry with MSR_EE")
> > Cc: stable@xxxxxxxxxxxxxxx # 6.8+
> > Reported-by: Timothy Pearson <tpearson@xxxxxxxxxxxxxxxxxxxxx>
> > Closes: https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.JavaMail.zimbra@xxxxxxxxxxxxxxxxxxxxxxxx
> > Signed-off-by: Gautam Menghani <gautam@xxxxxxxxxxxxx>
> > ---
> > v3:
> > 1. Continue the use of LPCR_MER when kernel-irqchip=off (Sashiko)
> >
> > v2:
> > 1. Handle the case where xive_interrupt_pending() is true and also the
> > external exception bit is set. (Narayana)
> >
> > arch/powerpc/include/asm/kvm_ppc.h | 7 +++++++
> > arch/powerpc/kvm/book3s_hv.c | 2 +-
> > 2 files changed, 8 insertions(+), 1 deletion(-)
> >
> > diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h
> > index 169ea6a7fbad..580ad2548c2b 100644
> > --- a/arch/powerpc/include/asm/kvm_ppc.h
> > +++ b/arch/powerpc/include/asm/kvm_ppc.h
> > @@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> > return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
> > }
> >
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
> > +{
> > + return kvm->arch.xive_devices.native;
> > +}
> > +
> > extern int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> > struct kvm_vcpu *vcpu, u32 cpu);
> > extern void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu);
> > @@ -782,6 +787,8 @@ static inline bool kvmppc_xive_rearm_escalation(struct kvm_vcpu *vcpu) { return
> >
> > static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> > { return 0; }
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm) { return false; }
> > +
> > static inline int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> > struct kvm_vcpu *vcpu, u32 cpu) { return -EBUSY; }
> > static inline void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu) { }
> > diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
> > index dbac3573b2c8..cb2bb29451a7 100644
> > --- a/arch/powerpc/kvm/book3s_hv.c
> > +++ b/arch/powerpc/kvm/book3s_hv.c
> > @@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 time_limit,
> > if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & MSR_EE))
> > kvmppc_inject_interrupt_hv(vcpu,
> > BOOK3S_INTERRUPT_EXTERNAL, 0);
> > - else
> > + else if (!kvmppc_xive_native_enabled(vcpu->kvm))
> > lpcr |= LPCR_MER;
>
> 1. What happens when an L1 KVM guest is booted with (native) XIVE and then the
> guest is rebooted with `xive=off` i.e., with XICS? Will we stop setting
> LPCR_MER as kvm->arch.xive_devices.native would still be set?

Yes, that's a good catch. kvm->arch.xive_devices.native is cleared only
when vm is destroyed. So this case has to be handled.

> 2. What happends when an L1 KVM guest is booted with XIVE and then we kexec into
> a new kernel with `xive=off`? You did mention in your other reply that with
> kexec, `xive=off` is ignored currently but IMO, we would want to understand
> where this limiation lies and fix that if need be. I do understand we are
> trying to fix spurious interrupts problem with XIVE in this patch and this
> particular problem can be taken separately but it worth investigating.

Yes right, the kexec case is to be understood, but I'll take it up separately.
Meanwhile, I'll send a v4 to fix the reboot case.

>
> Thanks,
> Amit