[PATCH v2 05/10] KVM: Destroy memslots immediately after mmu_notifiers are unregistered

From: Sean Christopherson

Date: Thu Oct 01 2026 - 16:24:49 EST


Instally dummy, empty memslots immediately after unregistering KVM's
mmu_notifier during VM destruction to harden against accessing memory from
the wrong address space when tearing down a VM. Because kvm_destroy_vm()
often runs when the associated VM's file is being released, current->mm is
often no longer kvm->mm, i.e. using the memslots to access userspace memory
is inherently broken/dangerous.

While KVM's APIs to read/write guest memory explicitly reject accesses if
current->mm != kvm->mm, taking away the memslots adds another layer of
defense and helps guard against rogue accesses that don't go through KVM's
standard API, or that do GUP+kmap().

Note, simply hoisting memslot destruction above kvm_arch_destroy_vm()
without installing dummy slots is not a viable alternative. Doing so would
require auditing all the paths of kvm_arch_destroy_vm(), which is nearly
infeasible, and missing even one case would result in a NULL pointer
dereference and/or use-after-free, neither of which is a substantially
better outcome than the status quo.

Signed-off-by: Sean Christopherson <seanjc@xxxxxxxxxx>
---
virt/kvm/kvm_main.c | 45 +++++++++++++++++++++++++++++++--------------
1 file changed, 31 insertions(+), 14 deletions(-)

diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index f368240aa1cd..770a2c3bd445 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -955,23 +955,42 @@ static void kvm_free_memslot(struct kvm *kvm, struct kvm_memory_slot *slot)
kfree(slot);
}

-static void kvm_free_memslots(struct kvm *kvm, struct kvm_memslots *slots)
+static const struct kvm_memslots kvm_empty_memslots = {
+ .generation = -1ull,
+ .hva_tree = RB_ROOT_CACHED,
+ .gfn_tree = RB_ROOT,
+ .id_hash[0 ... (ARRAY_SIZE(kvm_empty_memslots.id_hash) - 1)] = HLIST_HEAD_INIT,
+ .node_idx = 0,
+};
+
+static void kvm_destroy_memslots(struct kvm *kvm)
{
struct hlist_node *idnode;
struct kvm_memory_slot *memslot;
- int bkt;
+ int bkt, i;
+
+ /*
+ * Install empty memslots to guard against memslot lookups while the VM
+ * is being destroyed. Consuming memslots at this stage is a KVM bug,
+ * but "gracefully do nothing" is a much better outcome than "crash the
+ * host" or "corrupt random memory" when there inevitably is a bug.
+ */
+ mutex_lock(&kvm->slots_lock);
+ for (i = 0; i < kvm_arch_nr_memslot_as_ids(kvm); i++)
+ rcu_assign_pointer(kvm->memslots[i], &kvm_empty_memslots);
+
+ synchronize_srcu_expedited(&kvm->srcu);
+ mutex_unlock(&kvm->slots_lock);

/*
* The same memslot objects live in both active and inactive sets,
- * arbitrarily free using index '1' so the second invocation of this
- * function isn't operating over a structure with dangling pointers
- * (even though this function isn't actually touching them).
+ * arbitrarily free using index '1'.
*/
- if (!slots->node_idx)
- return;
-
- hash_for_each_safe(slots->id_hash, bkt, idnode, memslot, id_node[1])
- kvm_free_memslot(kvm, memslot);
+ for (i = 0; i < kvm_arch_nr_memslot_as_ids(kvm); i++) {
+ hash_for_each_safe(kvm->__memslots[i][1].id_hash, bkt, idnode,
+ memslot, id_node[1])
+ kvm_free_memslot(kvm, memslot);
+ }
}

static umode_t kvm_stats_debugfs_mode(const struct kvm_stats_desc *desc)
@@ -1302,12 +1321,10 @@ static void kvm_destroy_vm(struct kvm *kvm)
kvm->mn_active_invalidate_count = 0;
else
WARN_ON(kvm->mmu_invalidate_in_progress);
+ kvm_destroy_memslots(kvm);
+
kvm_arch_destroy_vm(kvm);
kvm_destroy_devices(kvm);
- for (i = 0; i < kvm_arch_nr_memslot_as_ids(kvm); i++) {
- kvm_free_memslots(kvm, &kvm->__memslots[i][0]);
- kvm_free_memslots(kvm, &kvm->__memslots[i][1]);
- }
cleanup_srcu_struct(&kvm->irq_srcu);
srcu_barrier(&kvm->srcu);
cleanup_srcu_struct(&kvm->srcu);
--
2.56.0.rc1.315.gc6ed9934b7-goog