Re: [PATCH 2/2] KVM: nVMX: avoid losing the posted notification interrupt when exiting L2

From: mlevitsk@xxxxxxxxxx

Date: Sat Oct 10 2026 - 13:03:29 EST


On Fri, 2026-10-09 at 16:50 -0400, Maxim Levitsky wrote:
> While delivering a nested posted interrupt notification
> (see vmx_deliver_nested_posted_interrupt), KVM assumes that,
> as long as the target vCPU is in the guest mode, KVM can send the
> special POSTED_INTR_NESTED_VECTOR, which will either trigger APICv ucode
> to inject all interrupts into L2 or cause a VM exit, after which
> vmx_complete_nested_posted_interrupt is supposed to finish the job.
>
> However, a third case is possible: if the target vCPU is about to exit
> to L1, the posted notification interrupt must be instead injected
> to L1' APIC, but KVM doesn't do this.
>
> Detect this case in the nested VM exit path and act accordingly.
>
> Signed-off-by: Maxim Levitsky <mlevitsk@xxxxxxxxxx>
> ---
>  arch/x86/kvm/vmx/nested.c | 10 +++++++++-
>  1 file changed, 9 insertions(+), 1 deletion(-)
>
> diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c
> index 2142e25c9b6c..4d8bb7f7fcd4 100644
> --- a/arch/x86/kvm/vmx/nested.c
> +++ b/arch/x86/kvm/vmx/nested.c
> @@ -2445,7 +2445,7 @@ static void prepare_vmcs02_early(struct vcpu_vmx *vmx, struct loaded_vmcs *vmcs0
>   ~PIN_BASED_VMX_PREEMPTION_TIMER);
>  
>   /* Posted interrupts setting is only taken from vmcs12.  */
> - vmx->nested.pi_pending = false;
> + WARN_ON_ONCE(vmx->nested.pi_pending);
>   if (nested_cpu_has_posted_intr(vmcs12)) {
>   vmx->nested.posted_intr_nv = vmcs12->posted_intr_nv;
>   } else {
> @@ -5162,6 +5162,14 @@ void __nested_vmx_vmexit(struct kvm_vcpu *vcpu, u32 vm_exit_reason,
>   kvm_clear_exception_queue(vcpu);
>   kvm_clear_interrupt_queue(vcpu);
>  
> + if (lapic_in_kernel(vcpu) && vmx->nested.pi_pending) {
> + vmx->nested.pi_pending = 0;
> + if (!WARN_ON_ONCE(vmx->nested.posted_intr_nv == -1)) {
> + kvm_lapic_set_irr(vmx->nested.posted_intr_nv, vcpu->arch.apic);
> + kvm_make_request(KVM_REQ_EVENT, vcpu);
> + }
> + }
> +

Hi!

After giving this a bit more of thought, I see that this will introduce spurious interrupts to L1, 
if the posted interrupt was actually served by APICv.

ON can't be trusted, so I only can add scan of PIR here to reduce chances of this happening, but I am not sure
that this can be completely avoided. 

I think that a spurious interrupt is better though that no interrupt because the guest can depend on it, and might
not even enter the L2, until it receives this posted interrupt.

What do you think?

Best regards,
Maxim Levitsky



>   vmx_switch_vmcs(vcpu, &vmx->vmcs01);
>  
>   kvm_nested_vmexit_handle_ibrs(vcpu);