Re: [PATCH v2] x86/kvm: introduce pv idle time

From: Sean Christopherson

Date: Fri Oct 02 2026 - 19:17:39 EST


+Vineeth and Josh

On Tue, Jul 21, 2026, He Rongguang wrote:
> Hi, this patch introduces a PV mechanism for guests to publish their
> vCPU idle state and accumulated idle time to the host. This allows
> the host to efficiently determine whether a vCPU is currently in its
> idle loop, and knows vCPU idled for how long. QEMU patch and ARM64
> support patch will follow in a subsequent series.
>
> The guest writes a GPA pointing to a struct kvm_idle_time via
> MSR_KVM_PV_IDLE_TIME. When entering idle, the guest sets the flag field
> to KVM_PV_VCPU_IDLE. On idle exit, it clears the flag back to
> KVM_PV_VCPU_RUNNING and adds the elapsed idle duration to idle_accum
> (in nanoseconds).
>
> The host can read the flag at any time through
> kvm_arch_is_vcpu_pv_idle() to make better scheduling or resource
> allocation decisions. For example, host may overcommit those vCPUs
> which are mostly idle. An in-guest agent may be absent, or may not
> report the status in time.
>
> The accumulated idle time provides visibility into per-vCPU usage for
> monitoring purposes. An in-guest agent can report VM CPU usage, but
> this PV mechanism allows the host to obtain guest CPU usage even when
> no agent is installed. This is especially helpful when the hypervisor
> enables exitless-hlt/mwait (for better guest performance) or when the
> guest uses idle halt-polling, both of which make QEMU vCPU thread usage
> deviate from the actual in-guest vCPU usage.
>
> This feature is advertised via KVM_FEATURE_PV_IDLE_TIME in CPUID leaf
> 0x40000001. Guests enable it by writing the appropriate MSR during
> initialization, similar to the existing steal time mechanism.
>
> Signed-off-by: He Rongguang <herongguang@xxxxxxxxxxxxxxxxx>
> ---
> v2:
> - grab kvm->srcu in kvm_arch_is_vcpu_pv_idle() before calling into
> kvm_read_guest_offset_cached().
> - return KVM_MSR_RET_UNSUPPORTED instead of 1 in MSR handling if
> !guest_pv_has(feature).
> - in set MSR handling, set vcpu->arch.pv_idle_time.msr_val after
> kvm_gfn_to_hva_cache_init() return success.
> - in kvm_arch_is_vcpu_pv_idle(), remove ghc->memslot check, let
> kvm_read_guest_offset_cached() handle it.
> - in kvm_arch_is_vcpu_pv_idle(), no need to use struct kvm_idle_time
> local variable, it is too big, just use __u64 flag is enough.
> - remove some straightforward code comments.
>
> v1:
> -
> https://lore.kernel.org/kvm/36863cf7-61b9-4885-946d-1608179fcab4@xxxxxxxxxxxxxxxxx/T/#u
> ---
> arch/x86/include/asm/kvm_host.h | 7 +++
> arch/x86/include/asm/kvm_para.h | 19 ++++++++
> arch/x86/include/uapi/asm/kvm_para.h | 24 ++++++++++
> arch/x86/kernel/kvm.c | 68 +++++++++++++++++++++++++++
> arch/x86/kernel/process.c | 7 +++
> arch/x86/kvm/cpuid.c | 3 +-
> arch/x86/kvm/x86.c | 69 ++++++++++++++++++++++++++++
> 7 files changed, 196 insertions(+), 1 deletion(-)

I really, really, reaaaaally don't want to take on any (more) PV scheduling ABI
in KVM, if possible. And I definitely don't want to take a bunch of one-off
hooks, e.g. for idle tracking and then for something else a few months/years
later.

Vineeth and Josh are working on PV scheduling via sched_ext and I assume additional
communication channels. I'm guessing idle state is one of the things that'll get
communicated to the host? I don't know the exact status of their work, but they've
got a slot in the sched_ext MC at LPC:

https://lpc.events/event/20/contributions/2482