Re: [PATCH v2 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 teardown race
From: Marc Zyngier
Date: Mon Aug 24 2026 - 11:56:07 EST
On Mon, 24 Aug 2026 14:15:29 +0100,
"Lorenzo Stoakes (ARM)" <ljs@xxxxxxxxxx> wrote:
>
> On Sun, Aug 23, 2026 at 08:53:31AM +0100, Marc Zyngier wrote:
> > On Sat, 22 Aug 2026 18:46:52 +0100,
> > "Lorenzo Stoakes (ARM)" <ljs@xxxxxxxxxx> wrote:
> > >
> > > When GFNs are invalidated in L0 an MMU notifier triggers
> > > kvm_unmap_gfn_range() which tears down all of the stage 2 shadow page
> > > tables for nested guests via kvm_nested_s2_unmap().
> > >
> > > To avoid lockup, the kvm->mmu_lock is dropped while doing this and the task
> > > rescheduled once for each block of physical address space (32 MiB for 16
> > > KiB page size), with the lock being reacquired once the task is scheduled
> > > again.
> > >
> > > This results in a potential race between this L0 tear down and tear down of
> > > the guest itself in kvm_flush_shadow_all(), a race which has been observed
> > > on real hardware.
> > >
> > > When this race occurs it causes an invalid kernel warning when the PGT of a
> > > nested MMU is cleared by kvm_flush_shadow_all() ->
> > > kvm_arch_flush_shadow_all() -> kvm_free_stage2_pgd().
> > >
> > > Patch 1 fixes this by having stage2_apply_range() no longer return an error
> > > when it has experienced a benign race with pgt teardown when it drops the
> > > lock.
> > >
> > > Patch 2 addresses something more serious - bad timing can turn this spurious
> > > warning into a NULL pointer dereference.
> > >
> > > kvm_arch_flush_shadow_all() calls kvm_uninit_stage2_mmu() which calls
> > > kvm_free_stage2_pgd() on the canonical kvm->arch.mmu for that guest's S2
> > > mappings, making it NULL.
> > >
> > > This is problematic if it happens before stage2_apply_range() reacquires
> > > the kvm->mmu_lock, as it ultimately returns to kvm_nested_s2_unmap() which
> > > dereferences kvm->arch.mmu.pgt with the mmu lock held under the incorrect
> > > assumption that it means it's valid, resulting in a NULL pointer
> > > dereference.
> > >
> > > Fix that by checking if kvm->arch.mmu.pgt is NULL before dereferencing it
> > > in kvm_nested_s2_unmap() and kvm_nested_s2_wp().
> >
> > With the commit message for patch #1 trimmed to something that fits on
> > a couple of standard terminal screens ;-) :
>
> Haha sure will put it on a diet and respin :)
Oliver can probably do so when applying the series.
Thanks,
M.
--
Jazz isn't dead. It just smells funny.