Re: [PATCH v2 0/3] KVM: TDX: Syntehsize SHUTDOWN on unhandled EPT Violation
From: Yan Zhao
Date: Thu Oct 08 2026 - 04:05:06 EST
On Tue, Sep 29, 2026 at 05:55:11PM -0700, Sean Christopherson wrote:
> On Wed, Sep 30, 2026, Rick P Edgecombe wrote:
> > On Tue, 2026-09-29 at 17:11 -0700, Sean Christopherson wrote:
> > > Signal SHUTDOWN instead of -EIO if the guest accesses an unaccepted page and
> > > has disabled EPT Violation #VEs on such accesses. Returning -EIO is all but
> > > guaranteed to mislead the VMM into thinking KVM (or the VMM) messed up, and
> > > will likely result in the VM being terminated instead of rebooted.
Hmm, it's by design in my previous implementation to terminate a VM when the
guest accesses an unaccepted page and has disabled EPT Violation #VEs on such
accesses, according to the TDX module spec:
"This happens if the TD is configured to TD-exit (instead of a #VE) on an EPT
violation due to accessing a PENDING page. It normally indicates an error
condition; the host VMM may decide to tear the TD down."
So, for the initial implementation, we chose to invoke kvm_vm_dead() and print
out the exact reason as a hint to the system admin. (Previously, reboot was also
not supported for a TDX guest).
Returning -EIO was based on the following considerations:
- vcpu_enter_guest() returns -EIO when kvm_test_request(KVM_REQ_VM_DEAD, vcpu)
is true.
- A request to handle a specific EPT violation is rejected by the firmware (the
TDX module).
> > I'm pretty sure Yan had a test for this path, and I'd love to see it actually
> > exercised. What is the urgency on getting this fix upstream? Can we wait a week?
I triggered the EPT violation in the kexec path, and successfully had the TD
- killed with error msg: "qemu-system-x86_64: cpus are not resettable,
terminating" on an old QEMU, or
- rebooted on a new QEMU with the msg printed:
"qemu-system-x86_64: info: virtual machine state has been rebuilt with new
guest file handle".
The TDX selftests we are using do not handle KVM_EXIT_SHUTDOWN, so if such
an error occurs, it's silently ignored by the TDX selftests even with msg
"kvm_intel: Guest access before accepting 0x8000c000 on vCPU 1" printed in dmesg.
> Absolutely. It can probably wait a month and no one would care. IIRC, this got
> hit by someone (internal to Google) deliberately crashing a guest kernel and doing
> funky things with kexec. It showed up on my radar purely because our automated
> madness alerted on the resulting assertion (on -EIO) in the VMM.
Though I didn't realize that returning -EIO would mislead the VMM into thinking
KVM (or the VMM) messed up, such an error could also be introduced by a VMM bug?
e.g., VMM removes an accepted page without any notification to the TD configured
with TDX_TD_ATTRIBUTES_SEPT_VE_DISABLE bit.