[PATCH v2 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 teardown race

From: Lorenzo Stoakes (ARM)

Date: Sat Aug 22 2026 - 13:47:22 EST


When GFNs are invalidated in L0 an MMU notifier triggers
kvm_unmap_gfn_range() which tears down all of the stage 2 shadow page
tables for nested guests via kvm_nested_s2_unmap().

To avoid lockup, the kvm->mmu_lock is dropped while doing this and the task
rescheduled once for each block of physical address space (32 MiB for 16
KiB page size), with the lock being reacquired once the task is scheduled
again.

This results in a potential race between this L0 tear down and tear down of
the guest itself in kvm_flush_shadow_all(), a race which has been observed
on real hardware.

When this race occurs it causes an invalid kernel warning when the PGT of a
nested MMU is cleared by kvm_flush_shadow_all() ->
kvm_arch_flush_shadow_all() -> kvm_free_stage2_pgd().

Patch 1 fixes this by having stage2_apply_range() no longer return an error
when it has experienced a benign race with pgt teardown when it drops the
lock.

Patch 2 addresses something more serious - bad timing can turn this spurious
warning into a NULL pointer dereference.

kvm_arch_flush_shadow_all() calls kvm_uninit_stage2_mmu() which calls
kvm_free_stage2_pgd() on the canonical kvm->arch.mmu for that guest's S2
mappings, making it NULL.

This is problematic if it happens before stage2_apply_range() reacquires
the kvm->mmu_lock, as it ultimately returns to kvm_nested_s2_unmap() which
dereferences kvm->arch.mmu.pgt with the mmu lock held under the incorrect
assumption that it means it's valid, resulting in a NULL pointer
dereference.

Fix that by checking if kvm->arch.mmu.pgt is NULL before dereferencing it
in kvm_nested_s2_unmap() and kvm_nested_s2_wp().

v2:
* Rebased onto next
* Updated 1/2's commit message to say that it was all of the kvmtool hosts
that were stopped, as per discussion with Wei-Lin and Yao Yuan.
* Updated 2/2 to remove the !may_block WARN_ON() as duplicative, as per
Marc.
* Updated 2/2 to abstract the VNCR IPA invalidation in
kvm_invalidate_vncr_ipa_all() and perform the same check for
kvm_nested_s2_wp(), as per discussion with Marc and sashiko report.

v1:
https://lore.kernel.org/r/20260812-kvm-arm-nested-virt-fix-v1-0-4ad883f1b6a5@xxxxxxxxxx

Signed-off-by: Lorenzo Stoakes (ARM) <ljs@xxxxxxxxxx>
---
Lorenzo Stoakes (ARM) (2):
KVM: arm64: Fix spurious warning for benign stage 2 teardown race
KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race

arch/arm64/kvm/mmu.c | 10 ++++++++--
arch/arm64/kvm/nested.c | 15 +++++++++++++--
2 files changed, 21 insertions(+), 4 deletions(-)
---
base-commit: aa8e5dc6a7a2a1141ab40706a51010adcd0e57d2
change-id: 20260811-kvm-arm-nested-virt-fix-031e9ab1be87

Best regards,
--
Lorenzo Stoakes (ARM) <ljs@xxxxxxxxxx>