Re: [PATCH v2] KVM: nVMX: Service local TLB flushes on failed nested VM-Enter

From: Yosry Ahmed

Date: Mon Jul 27 2026 - 20:44:01 EST


On Wed, Jul 22, 2026 at 11:01:28PM +0000, Yosry Ahmed wrote:
> KVM services local TLB flushes on "full" nested VM-Exits (through
> __nested_vmx_vmexit()), but not if a nested VM-Enter fails (e.g. due to
> failed VMCS checks in nested_vmx_enter_non_root_mode()).
>
> However, it is possible that KVM had queued TLB flushes that need to be
> performed, even if the nested VM-Enter was not successful. For example,
> if VPID is disabled for L2 (via nested_vmx_transition_tlb_flush(), or if
> via the MSR load lists, as the SDM says:
>
> If any MSR is being loaded in such a way that would architecturally
> require a TLB flush, the TLBs are updated so that, after VM entry, the
> logical processor will not use any translations that were cached before
> the transition.
>
> The SDM is unclear about when the TLB flush should occur, and whether or
> not a failed VM entry would flush the TLB, so it is safer to always
> do the TLB flush in this case.
>
> More concretely, KVM also updates the last VPID L1 used for L2 in
> nested_vmx_transition_tlb_flush() (i.e. last_vpid), even if the VM entry
> ultimately fails. With the current code, KVM could miss a TLB flush if
> L1 changes L2's VPID, then does a failed VM entry followed by a
> successful one, as the failed VM entry would update last_vpid but not
> actually flush the TLB. Servicing local TLB flushes on failed VM entries
> makes sure that the TLB is always flushed when last_vpid is updated.
>
> Fixes: 5c614b3583e7 ("KVM: nVMX: nested VPID emulation")
> Cc: stable@xxxxxxxxxxxxxxx
> Reported-by: Sashiko <sashiko-bot@xxxxxxxxxx> # Internal review
> Suggested-by: Sean Christopherson <seanjc@xxxxxxxxxx>
> Signed-off-by: Yosry Ahmed <yosry@xxxxxxxxxx>
> ---

For the record, my reproducer was basically the nested TLB flushes
selftest introduced here:
https://lore.kernel.org/kvm/20260728003557.1136583-29-yosry@xxxxxxxxxx/

With this diff on top:

diff --git a/tools/testing/selftests/kvm/x86/nested_tlb_flush_test.c b/tools/testing/selftests/kvm/x86/nested_tlb_flush_test.c
index 55c9909bb085f..1659dd9c0044a 100644
--- a/tools/testing/selftests/kvm/x86/nested_tlb_flush_test.c
+++ b/tools/testing/selftests/kvm/x86/nested_tlb_flush_test.c
@@ -123,8 +123,26 @@ static void run_l2(void *nested_state, bool launch)
vmx_tlb_flush(nested_state);
if (launch)
GUEST_ASSERT(!vmlaunch());
- else
- GUEST_ASSERT(!vmresume());
+ else {
+ struct vmx_pages *vmx = nested_state;
+ struct vmx_msr_entry *entry = vmx->msr;
+
+ /* Inject VM-Entry failure via invalid MSR load */
+ entry->index = 0xc0000100; /* MSR_FS_BASE, disallowed for loading */
+ entry->reserved = 0;
+ entry->value = 0;
+ vmwrite(VM_ENTRY_MSR_LOAD_ADDR, vmx->msr_gpa);
+ vmwrite(VM_ENTRY_MSR_LOAD_COUNT, 1);
+
+ GUEST_ASSERT_EQ(vmresume(), 0);
+ GUEST_ASSERT_EQ(vmreadz(VM_EXIT_REASON), (EXIT_REASON_FAILED_VMENTRY | EXIT_REASON_MSR_LOAD_FAIL));
+
+ /* Fix failure and retry */
+ vmwrite(VM_ENTRY_MSR_LOAD_COUNT, 0);
+ memset(vmx->msr, 0, 4096);
+
+ GUEST_ASSERT_EQ(vmresume(), 0);
+ }
GUEST_ASSERT_EQ(vmreadz(VM_EXIT_REASON), EXIT_REASON_VMCALL);
vmwrite(GUEST_RIP, vmreadz(GUEST_RIP) + 3); /* skip over VMCALL */
} else {