[PATCH] x86/kvm: Don't consume async #PF flags on a kernel #PF in NMI context
From: EJ Campbell
Date: Tue Sep 29 2026 - 14:14:43 EST
An async #PF "page not present" event arrives through the #PF vector,
and its reason sits in the per-CPU apf_reason.flags until
__kvm_handle_async_pf() reads and clears it. If an NMI arrives after
the event is injected and before the #PF handler reads the flags, and
the NMI handler itself takes a #PF, that fault reads and clears the
flags belonging to the #PF it interrupted. It then finds interrupts
disabled in the NMI context it interrupted, and panics:
Kernel panic - not syncing: Host injected async #PF in interrupt
disabled region
Call Trace:
<NMI>
panic+0x52/0x60
__kvm_handle_async_pf+0xb0/0xb0
exc_page_fault+0x71/0xb0
asm_exc_page_fault+0x27/0x30
RIP: 0010:__get_user_nocheck_8+0x6/0x20
perf_callchain_user+0x147/0x1a0
get_perf_callchain+0x12d/0x1c0
...
perf_event_nmi_handler+0x26/0x50
exc_nmi+0xe8/0x110
RIP: 0010:asm_exc_page_fault+0x0/0x30
</NMI>
The last RIP line shows the NMI landing on the first instruction of
the #PF entry. perf with user call chains takes such a fault: the
sampling NMI walks the user stack with __get_user(), which faults when
a saved frame pointer is not a valid user address or the walk runs off
the top of the stack. The same window exists when KVM, running as an
L1 hypervisor with DELIVERY_AS_PF_VMEXIT, reads the flags itself after
a #PF VM-exit.
The NMI entry code saves and restores CR2 for this nesting, but not the
async #PF reason. Andy Lutomirski described the race, with perf as the
example, when async #PF delivery was reworked in 2020 [1]. Saving and
clearing the flags across the NMI, next to CR2, was discussed there but
not merged.
Since that rework the guest does not set KVM_ASYNC_PF_SEND_ALWAYS, so
the host delivers "page not present" only to user mode. Read the flags
only for a fault from user mode. A kernel-mode fault in NMI context
(in_nmi() is also true in kernel-mode #MC and #DB) returns without
touching the flags, and the interrupted handler finds them intact. This
leaves the NMI entry path unchanged. A kernel-mode fault outside NMI
context only checks whether a reason is pending, and keeps the "Host
injected async #PF in kernel mode" panic for a host that delivers one.
An earlier posting [2] also skipped the flags for kernel-mode faults,
for a different reason, and removed that panic.
Tested on v7.3-rc5 in a 4-vCPU, 2 GiB guest restored from a snapshot
whose memory is served on demand through userfaultfd, so that reading
it back takes async #PFs, while every thread samples hardware cycles
with user call chains and points its frame pointer at an unmapped page.
Without this patch the guest panicked as above in 3 of 3 runs. With it,
the guest survived 3 of 3 runs of 60 seconds, and the host counted
148,441 to 153,334 kvm_async_pf_not_present events in the first 20
seconds of each. On 6.18.50 the results were the same: the guest
panicked without the patch and survived with it.
Link: https://lore.kernel.org/all/ed71d0967113a35f670a9625a058b8e6e0b2f104.1583547991.git.luto@xxxxxxxxxx/
[1]
Link: https://lore.kernel.org/all/20211126123145.2772-1-jiangshanlai@xxxxxxxxx/
[2]
Cc: stable@xxxxxxxxxxxxxxx # v5.9+
Signed-off-by: EJ Campbell <ejc3@xxxxxxxx>
---
Notes:
- The reproducer is a small C program run in a guest restored by fcvm,
a Firecracker-based VM manager:
https://github.com/ejc3/fcvm/blob/9c7519f861273e8be954f9d006bc4d6654669f15/tests/data/async_pf_nmi.c
It fills memory, the guest is snapshotted, and the restored guest reads
the memory back while it samples with call chains.
- For a kernel-mode fault outside NMI context with a reason pending,
which requires a host that ignores KVM_ASYNC_PF_SEND_ALWAYS being
clear, two messages change: with IF=0 the guest now panics with "Host
injected async #PF in kernel mode" rather than "... in interrupt
disabled region", and a reason other than PAGE_NOT_PRESENT panics
instead of hitting the WARN_ONCE. The guest requires ASYNC_PF_INT, so
PAGE_NOT_PRESENT is the only reason a conforming host writes.
- No Fixes: tag. The race predates ef68017eb570 ("x86/kvm: Handle async
page faults directly through do_page_fault()"), which created
__kvm_handle_async_pf(). The fix relies on the host not delivering
"page not present" to kernel mode (since v5.8) and on "page ready"
arriving by interrupt rather than #PF (since v5.9), hence v5.9+.
---
arch/x86/kernel/kvm.c | 30 +++++++++++++++++++++++++++---
1 file changed, 27 insertions(+), 3 deletions(-)
diff --git a/arch/x86/kernel/kvm.c b/arch/x86/kernel/kvm.c
index 6b0a5861c..87ba4e174 100644
--- a/arch/x86/kernel/kvm.c
+++ b/arch/x86/kernel/kvm.c
@@ -265,11 +265,37 @@ noinstr u32 kvm_read_and_reset_apf_flags(void)
}
EXPORT_SYMBOL_FOR_KVM(kvm_read_and_reset_apf_flags);
+static __always_inline bool kvm_apf_reason_pending(void)
+{
+ return __this_cpu_read(async_pf_enabled) &&
+ __this_cpu_read(apf_reason.flags);
+}
+
noinstr bool __kvm_handle_async_pf(struct pt_regs *regs, u32 token)
{
- u32 flags = kvm_read_and_reset_apf_flags();
irqentry_state_t state;
+ u32 flags;
+
+ /*
+ * The guest does not enable KVM_ASYNC_PF_SEND_ALWAYS, so the host
+ * delivers async #PF only to user mode (or as a #PF VM-exit, which KVM
+ * reads itself). A #PF from kernel mode can still arrive between that
+ * delivery and the read of the reason flags, inside an NMI that
+ * interrupted it, for instance perf reading user memory for a call
+ * chain. Such a fault is a real one: leave the flags for the fault they
+ * belong to. Outside NMI context, which in_nmi() also reports for
+ * kernel-mode #MC and #DB, pending flags on a kernel-mode fault mean
+ * the host is broken.
+ */
+ if (!user_mode(regs)) {
+ if (likely(in_nmi() || !kvm_apf_reason_pending()))
+ return false;
+ state = irqentry_enter(regs);
+ instrumentation_begin();
+ panic("Host injected async #PF in kernel mode\n");
+ }
+ flags = kvm_read_and_reset_apf_flags();
if (!flags)
return false;
@@ -285,8 +311,6 @@ noinstr bool __kvm_handle_async_pf(struct pt_regs
*regs, u32 token)
panic("Host injected async #PF in interrupt disabled region\n");
if (flags & KVM_PV_REASON_PAGE_NOT_PRESENT) {
- if (unlikely(!(user_mode(regs))))
- panic("Host injected async #PF in kernel mode\n");
/* Page is swapped out by the host. */
kvm_async_pf_task_wait_schedule(token);
} else {
---
base-commit: 6f8319e3e9a44dd537d17f41565a8453c560a581
change-id: 20260929-kvm-apf-nmi-1911965f5e46
Best regards,
--
EJ Campbell <ejc3@xxxxxxxx>