Re: [PATCH 3/4] KVM: x86/mmu: Bug the VM if KVM calcs a CPU role with EFER.LMA=1 && CR4.PAE=0

From: Sean Christopherson

Date: Thu Aug 27 2026 - 13:30:37 EST


On Thu, Aug 27, 2026, Yosry Ahmed wrote:
> On Thu, Aug 27, 2026 at 7:57 AM Sean Christopherson <seanjc@xxxxxxxxxx> wrote:
> > > Can we shove this into the existing if (____is_efer_lma(regs)) below?
> >
> > No, because there are three more checks on EFER.LMA:
> >
> > role.ext.cr4_pke = ____is_efer_lma(regs) && ____is_cr4_pke(regs);
> > role.ext.cr4_la57 = ____is_efer_lma(regs) && ____is_cr4_la57(regs);
> > role.ext.efer_lma = ____is_efer_lma(regs);
> >
> > and I don't want to have to condition them all on something that shouldn't happen.
>
> Yeah I assumed that we don't care about the state anymore if we'll
> KVM_BUG_ON(), but apparently that's not the case based on your comment below.

Ya, it's not an immediate "jump all the way back to userspace", though that would
be kinda cool/terrifying.

> > OMG, I hate SVM. I resurrected the selftest hack I used to verify this bug, to
> > demonstrate that Sashiko's "technically that's undefined behavior and this is
> > useless" complaint is wrong, because even though it's undefined behavior and the
> > compiler *could* ignore the change, in practice the compiler probably won't ignore
> > the change. And since this is defense-in-depth, it's "fine" if the paranoid
> > hardening only isn't guaranteed to kick in.
> >
> > And in doing so managed to trip this KVM_BUG_ON() in *L0* when running the test
> > in L1, because as you kinda sorta noted in patch 1, KVM doesn't ignore EFER.LMA
> > when loading L2 state.
> >
> > I had actually tried to do exactly that, by having nested_vmcb_check_save() clear
> > EFER.LMA if EFER.LME=0, but that doesn't work because svm_set_nested_state() uses
> > the "cache" only for the checks, not for the actual loading of state. *sigh*
> >
> > So in addition to patch 1, we also need this to guard against configuring L2's
> > walk_mmu with bad state.
>
> Hmm wouldn't it be simpler at this point to let KVM_SET_NESTED_STATE
> and nested VMRUN have the invalid LMA/LME combination and just ignore
> EFER.LMA if EFER.LME

Definitely not a straight "ignore", because that would end up being an even worse
game of whack-a-mole, because very path that checks vcpu->arch.efer would have to
account for that possibility.

We could forcefully sanitize EFER in flows that write EFER, but (a) that's still
a (must smaller) game of whack-a-mole and (b) it would actively hide KVM bugs for
flows that are supposed to reject the invalid state. And if we WARNed to address
(b), we'll be right back where we are today: playing whack-a-mole to prevent the
WARN from being triggered.

> (or just always check EFER.LMA && EFER.LME)?

No can do, because we can't disallow the combination for L2 on VMRUN without
violating AMD's architecture. And practically speaking, we *are* doing that,
just in a bunch of places because there's no one rule to rule them all.

> > Because there's a lot of code between here and checking KVM_VM_DEAD in
> > vcpu_enter_guest(). And has been proven far too many times this year, detecting
> > a flaw doesn't automagically mitigate true badness.
>
> Interesting, I always assumed we can do whatever we want after KVM_BUG_ON() :P

Nope. In addition to KVM_VM_DEAD not being checked until vcpu_enter_guest(),
more broadly it only kicks in for cross-task behaviors on the next ioctl. E.g.
if KVM_BUG_ON() guards against bad VM state, as opposed to bad vCPU state, i.e.
if *other* vCPUs could consume the bad state, then it's especially important to
take evasive action.

KVM_BUG_ON() is as much about protecting the guest as it is about protecting the
host. E.g. if KVM *knows* it fatally screwed up, then continuing to run the
guest risks corrupting guest state and thus causing far worse problems than DoSing
the guest.