[PATCH v2 08/10] KVM: nVMX: Don't flush shadow VMCS12 to guest memory during vCPU teardown

From: Sean Christopherson

Date: Thu Oct 01 2026 - 16:25:42 EST


From: Jim Mattson <jmattson@xxxxxxxxxx>

When a vCPU is destroyed while L2 is active, KVM synthesizes a nested
VM-Exit, which flushes the cached shadow VMCS12 back to guest memory:

vmx_vcpu_free()
|-> nested_vmx_free_vcpu()
|-> vmx_leave_nested()
|-> nested_vmx_vmexit(vcpu, -1, 0, 0)
|-> nested_flush_cached_shadow_vmcs12()
|-> kvm_write_guest_cached()
|-> __copy_to_user(ghc->hva, ...)

Accessing user memory via a memslot during VM destruction is broken, as
there are no guarantees that current->mm == kvm->mm when the VM is dying,
because the last reference to the VM can be put from a different process
than the original creating processes.

And even if the original process does put the final reference, during
process exit, do_exit() calls exit_mm() before closing file descriptors, so
vCPU destruction runs with current->mm == NULL on a borrowed lazy TLB
active_mm. If the borrowed address space has a writable mapping at the
to-be-written userspace address, KVM will corrupt an unrelated task's
memory since uaccess APIs, including __copy_to_user(), don't sanity check
current->mm (and *can't* sanity perform KVM's current->mm == kvm->mm check
since that is firmly a KVM-only concept).

Hack-a-fix the nVMX flow even though KVM now protects against bad uaccess
reads/writes in the core APIs, as doing so will allow adding even more
sanity checks in KVM's APIs to help detect other buggy code. Add a TODO to
call out that checking if KVM can do a uaccess for the VM is a hack; nVMX
really needs to stop abusing __nested_vmx_vmexit() when destroying a vCPU.

Fixes: 61ada7488ffd ("KVM: nVMX: Cache shadow vmcs12 on VMEntry and flush to memory on VMExit")
Signed-off-by: Jim Mattson <jmattson@xxxxxxxxxx>
[sean: key off __kvm_can_do_uaccess(), add TODO]
Signed-off-by: Sean Christopherson <seanjc@xxxxxxxxxx>
---
arch/x86/kvm/vmx/nested.c | 7 ++++++-
include/linux/kvm_host.h | 8 ++++++--
2 files changed, 12 insertions(+), 3 deletions(-)

diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c
index 151873407abd..8e31eba4d9fa 100644
--- a/arch/x86/kvm/vmx/nested.c
+++ b/arch/x86/kvm/vmx/nested.c
@@ -5134,8 +5134,13 @@ void __nested_vmx_vmexit(struct kvm_vcpu *vcpu, u32 vm_exit_reason,
* Otherwise, this flush will dirty guest memory at a
* point it is already assumed by user-space to be
* immutable.
+ *
+ * TODO: Drop the explicit check on being able to access guest
+ * memory once KVM no longer abuses the nested VM-Exit
+ * flow when destroying a vCPU.
*/
- nested_flush_cached_shadow_vmcs12(vcpu, vmcs12);
+ if (__kvm_can_do_uaccess(vcpu->kvm))
+ nested_flush_cached_shadow_vmcs12(vcpu, vmcs12);
} else {
/*
* The only expected VM-instruction error is "VM entry with
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 0ee81754d730..0c58a4945595 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -1350,10 +1350,14 @@ int kvm_write_guest_offset_cached(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
int kvm_gfn_to_hva_cache_init(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
gpa_t gpa, unsigned long len);

+static __always_inline __must_check bool __kvm_can_do_uaccess(struct kvm *kvm)
+{
+ return current->mm == kvm->mm && refcount_read(&kvm->users_count);
+}
+
static __always_inline __must_check bool kvm_can_do_uaccess(struct kvm *kvm)
{
- return !WARN_ON_ONCE(current->mm != kvm->mm ||
- !refcount_read(&kvm->users_count));
+ return !WARN_ON_ONCE(!__kvm_can_do_uaccess(kvm));
}

#define BUILD_KVM_COPY_USER_WRAPPER(fn, to_user, from_user) \
--
2.56.0.rc1.315.gc6ed9934b7-goog