[PATCH v3 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 teardown race

From: Lorenzo Stoakes (ARM)

Date: Tue Sep 01 2026 - 13:42:36 EST


When GFNs are invalidated in L0 an MMU notifier triggers
kvm_unmap_gfn_range() which tears down all of the stage 2 shadow page
tables for nested guests via kvm_nested_s2_unmap().

To avoid lockup, the kvm->mmu_lock is dropped while doing this and the task
rescheduled once for each block of physical address space (32 MiB for 16
KiB page size), with the lock being reacquired once the task is scheduled
again.

This results in a potential race between this L0 tear down and tear down of
the guest itself in kvm_flush_shadow_all(), a race which has been observed
on real hardware.

When this race occurs it causes an invalid kernel warning when the PGT of a
nested MMU is cleared by kvm_flush_shadow_all() ->
kvm_arch_flush_shadow_all() -> kvm_free_stage2_pgd().

Patch 1 fixes this by having stage2_apply_range() no longer return an error
when it has experienced a benign race with pgt teardown when it drops the
lock.

Patch 2 addresses something more serious - bad timing can turn this spurious
warning into a NULL pointer dereference.

kvm_arch_flush_shadow_all() calls kvm_uninit_stage2_mmu() which calls
kvm_free_stage2_pgd() on the canonical kvm->arch.mmu for that guest's S2
mappings, making it NULL.

This is problematic if it happens before stage2_apply_range() reacquires
the kvm->mmu_lock, as it ultimately returns to kvm_nested_s2_unmap() which
dereferences kvm->arch.mmu.pgt with the mmu lock held under the incorrect
assumption that it means it's valid, resulting in a NULL pointer
dereference.

Fix that by checking if kvm->arch.mmu.pgt is NULL before dereferencing it
in kvm_nested_s2_unmap() and kvm_nested_s2_wp().

v3:
* Rebased onto Linus's tree.
* Added R-b tags (thanks Yao and Marc!).
* Put commit message for 1/2 on a diet as requested by Marc.
* Clarified logic in stage2_apply_range() as per Yao Yuan.

v2:
* Rebased onto next
* Updated 1/2's commit message to say that it was all of the kvmtool hosts
that were stopped, as per discussion with Wei-Lin and Yao Yuan.
* Updated 2/2 to remove the !may_block WARN_ON() as duplicative, as per
Marc.
* Updated 2/2 to abstract the VNCR IPA invalidation in
kvm_invalidate_vncr_ipa_all() and perform the same check for
kvm_nested_s2_wp(), as per discussion with Marc and sashiko report.
https://lore.kernel.org/r/20260822-kvm-arm-nested-virt-fix-v2-0-ac4059a0eaa6@xxxxxxxxxx

v1:
https://lore.kernel.org/r/20260812-kvm-arm-nested-virt-fix-v1-0-4ad883f1b6a5@xxxxxxxxxx

Signed-off-by: Lorenzo Stoakes (ARM) <ljs@xxxxxxxxxx>
---
Lorenzo Stoakes (ARM) (2):
KVM: arm64: Fix spurious warning for benign stage 2 teardown race
KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race

arch/arm64/kvm/mmu.c | 15 ++++++++++++---
arch/arm64/kvm/nested.c | 15 +++++++++++++--
2 files changed, 25 insertions(+), 5 deletions(-)
---
base-commit: 786262be6048deab760f68c8acc2c85607165894
change-id: 20260811-kvm-arm-nested-virt-fix-031e9ab1be87

Best regards,
--
Lorenzo Stoakes (ARM) <ljs@xxxxxxxxxx>