Re: [PATCH v3 10/10] KVM: selftests: Trigger L2->L1 exits stress save+restore and #PF test

From: Yosry Ahmed

Date: Fri Jul 24 2026 - 14:18:55 EST


On Fri, Jul 24, 2026 at 10:50 AM Sean Christopherson <seanjc@xxxxxxxxxx> wrote:
>
> On Mon, Jun 29, 2026, Yosry Ahmed wrote:
> > Extend the testing coverage in L2 by injecting a #UD into the vCPU every
> > other iteration during restore, and intercepting #UD from L1,
> > essentially forcing an L2 -> L1 VM-Exit directly after save+restore.
>
> Assuming the #UD is a means to an end, make that the focus of the intro blurb.
> I read this changelog without looking at the shortlog, and was about to ask why
> injecting a #UD is interesting, and then I saw the comment. It'd be helpful
> to add a bit more context too, as it took me a few seconds to piece together
> that the goal is to force the exit while L0 has control, i.e. a more obvious
> hypercall from L2 wouldn't suffice.

Yes, #UD is the easiest way (I think) to force an exit to L1 without
running L2 first. Will reword.

>
> > With this change, the test reliably reproduces the CR2 bug fixed by
> > commit 5c247d08bc81 ("KVM: nSVM: Use vcpu->arch.cr2 when updating vmcb12
> > on nested #VMEXIT") -- at least on Milan, Genoa, and Turin CPUs.
> >
> > Assisted-by: Gemini:gemini-3.1-pro
> > Signed-off-by: Yosry Ahmed <yosry@xxxxxxxxxx>
> > ---
> > .../kvm/x86/stress_save_restore_pf_test.c | 47 +++++++++++++++++--
> > 1 file changed, 42 insertions(+), 5 deletions(-)
> >
> > diff --git a/tools/testing/selftests/kvm/x86/stress_save_restore_pf_test.c b/tools/testing/selftests/kvm/x86/stress_save_restore_pf_test.c
> > index 9ab52d27a61d9..2b76e56f744e7 100644
> > --- a/tools/testing/selftests/kvm/x86/stress_save_restore_pf_test.c
> > +++ b/tools/testing/selftests/kvm/x86/stress_save_restore_pf_test.c
> > @@ -105,8 +105,12 @@ static void guest_access_memory(void *arg)
> > static void l1_svm_code(struct svm_test_data *svm)
> > {
> > generic_svm_setup(svm, guest_access_memory);
> > - run_guest(svm->vmcb, svm->vmcb_gpa);
> > - GUEST_ASSERT(false);
> > + svm->vmcb->control.intercept_exceptions |= BIT(UD_VECTOR);
> > +
> > + while (1) {
> > + run_guest(svm->vmcb, svm->vmcb_gpa);
> > + GUEST_ASSERT_EQ(svm->vmcb->control.exit_code, (SVM_EXIT_EXCP_BASE + UD_VECTOR));
>
> Please wrap this one, it's so long that I find it genuinely difficult to parse.
>
> GUEST_ASSERT_EQ(svm->vmcb->control.exit_code,
> (SVM_EXIT_EXCP_BASE + UD_VECTOR));

Will do.

> > static void l1_guest_code(void *test_data)
> > @@ -159,6 +167,24 @@ static void vcpu_sigusr_ignore(void)
> > sigaction(SIGUSR1, &sa, NULL);
> > }
> >
> > +static bool vcpu_state_is_guest_mode(struct kvm_x86_state *state)
> > +{
> > + return !!(state->nested.flags & KVM_STATE_NESTED_GUEST_MODE);
> > +}
> > +
> > +static void vcpu_state_inject_ud(struct kvm_x86_state *state)
> > +{
> > + if (state->events.exception.pending || state->events.exception.injected)
> > + return;
> > +
> > + state->events.flags |= KVM_VCPUEVENT_VALID_PAYLOAD;
> > + state->events.exception.pending = true;
> > + state->events.exception.injected = false;
> > + state->events.exception.nr = UD_VECTOR;
> > + state->events.exception.has_error_code = false;
> > + state->events.exception_has_payload = false;
> > +}
> > +
> > static bool parse_args_nested(int argc, char *argv[])
> > {
> > bool nested = false;
> > @@ -192,10 +218,13 @@ int main(int argc, char *argv[])
> > gva_t gva;
> > u64 pte;
> >
> > + TEST_REQUIRE(kvm_has_cap(KVM_CAP_EXCEPTION_PAYLOAD));
>
> But KVM_CAP_EXCEPTION_PAYLOAD _isn't_ required, it's an optional feature. Actually,
> this is ridiculous. The test is injecting a #UD, it doesn't have a payload.
> Bad AI, bad.

KVM_CAP_EXCEPTION_PAYLOAD is required to inject a *pending* exception,
which is needed as KVM won't check for interception on already
injected exceptions. I will add a comment.

> > +
> > nested = parse_args_nested(argc, argv);
> >
> > vm = vm_create_with_one_vcpu(&vcpu, nested ? l1_guest_code : guest_access_memory);
> > vm_install_exception_handler(vm, PF_VECTOR, guest_pf_handler);
> > + vm_enable_cap(vm, KVM_CAP_EXCEPTION_PAYLOAD, -2ul);
>
> -2ul?

That was actually me, not the AI. Apparently the other selftests also
write -2ul for some reason. I was too lazy to go change them and/or
figure out why -2ul, it didn't make any sense to me. I chose the lazy
option and just used the same thing here.

Someone named Sean Christopherson wrote the other two, maybe we can ask him? :P

>
> > if (nested) {
> > TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_SVM) || kvm_cpu_has(X86_FEATURE_VMX));
> > @@ -270,8 +299,16 @@ int main(int argc, char *argv[])
> >
> > state = vcpu_save_state(vcpu);
> >
> > + /*
> > + * If the vCPU is in guest mode, inject a #UD to trigger an
> > + * L2->L1 VM-Exit every other iteration.
> > + */
> > + if (nested && vcpu_state_is_guest_mode(state) && count % 2 == 0)
>
> Checking "nested" here is unnecessary.

It is, actually. Sashiko pointed out [1] that
vcpu_state_is_guest_mode() will read uninitialized memory otherwise.

[1]https://sashiko.dev/#/patchset/20260518202514.2037078-1-yosry%40kernel.org?part=8

>
> > + vcpu_state_inject_ud(state);
>
> Honestly, I'd rather open code this whole thing, because this doesn't actually
> inject a #UD. It _prepares_ state, but doesn't send that into KVM. E.g.

It injects a #UD into the state. The function name and parameter
should make it clear. I prefer the helper, but I won't die on this
hill.

>
> /*
> * If the vCPU is in guest mode, inject a #UD to trigger an
> * L2->L1 VM-Exit every other iteration, unless the vCPU has. Take care not to
> * clobber any exceptions
> */
> if ((i & 1) && (state.nested.flags & KVM_STATE_NESTED_GUEST_MODE) &&
> !state.events.exception.pending && !state.events.exception.injected) {
> state->events.exception.pending = true;
> state->events.exception.injected = false;
> state->events.exception.nr = UD_VECTOR;
> state->events.exception.has_error_code = false;
> }
>
> > +
> > kvm_vm_release(vm);
> > vcpu = vm_recreate_with_one_vcpu(vm);
> > + vm_enable_cap(vm, KVM_CAP_EXCEPTION_PAYLOAD, -2ul);
> > vcpu_load_state(vcpu, state);
> > kvm_x86_state_cleanup(state);
> >
> > --
> > 2.55.0.rc0.799.gd6f94ed593-goog
> >