Re: [PATCH v2] x86/kvm: introduce pv idle time
From: Vineeth Remanan Pillai
Date: Mon Oct 05 2026 - 11:26:22 EST
On Mon, Oct 5, 2026 at 12:20 AM Josh Don <joshdon@xxxxxxxxxx> wrote:
>
> On Sat, Oct 3, 2026 at 1:17 AM Sean Christopherson <seanjc@xxxxxxxxxx> wrote:
> >
> > +Vineeth and Josh
> >
> > On Tue, Jul 21, 2026, He Rongguang wrote:
> > > Hi, this patch introduces a PV mechanism for guests to publish their
> > > vCPU idle state and accumulated idle time to the host. This allows
> > > the host to efficiently determine whether a vCPU is currently in its
> > > idle loop, and knows vCPU idled for how long. QEMU patch and ARM64
> > > support patch will follow in a subsequent series.
> > >
> > > The guest writes a GPA pointing to a struct kvm_idle_time via
> > > MSR_KVM_PV_IDLE_TIME. When entering idle, the guest sets the flag field
> > > to KVM_PV_VCPU_IDLE. On idle exit, it clears the flag back to
> > > KVM_PV_VCPU_RUNNING and adds the elapsed idle duration to idle_accum
> > > (in nanoseconds).
> > >
> > > The host can read the flag at any time through
> > > kvm_arch_is_vcpu_pv_idle() to make better scheduling or resource
> > > allocation decisions. For example, host may overcommit those vCPUs
> > > which are mostly idle. An in-guest agent may be absent, or may not
> > > report the status in time.
> > >
> > > The accumulated idle time provides visibility into per-vCPU usage for
> > > monitoring purposes. An in-guest agent can report VM CPU usage, but
> > > this PV mechanism allows the host to obtain guest CPU usage even when
> > > no agent is installed. This is especially helpful when the hypervisor
> > > enables exitless-hlt/mwait (for better guest performance) or when the
> > > guest uses idle halt-polling, both of which make QEMU vCPU thread usage
> > > deviate from the actual in-guest vCPU usage.
> > >
> > > This feature is advertised via KVM_FEATURE_PV_IDLE_TIME in CPUID leaf
> > > 0x40000001. Guests enable it by writing the appropriate MSR during
> > > initialization, similar to the existing steal time mechanism.
> > >
> > > Signed-off-by: He Rongguang <herongguang@xxxxxxxxxxxxxxxxx>
> > > ---
> > > v2:
> > > - grab kvm->srcu in kvm_arch_is_vcpu_pv_idle() before calling into
> > > kvm_read_guest_offset_cached().
> > > - return KVM_MSR_RET_UNSUPPORTED instead of 1 in MSR handling if
> > > !guest_pv_has(feature).
> > > - in set MSR handling, set vcpu->arch.pv_idle_time.msr_val after
> > > kvm_gfn_to_hva_cache_init() return success.
> > > - in kvm_arch_is_vcpu_pv_idle(), remove ghc->memslot check, let
> > > kvm_read_guest_offset_cached() handle it.
> > > - in kvm_arch_is_vcpu_pv_idle(), no need to use struct kvm_idle_time
> > > local variable, it is too big, just use __u64 flag is enough.
> > > - remove some straightforward code comments.
> > >
> > > v1:
> > > -
> > > https://lore.kernel.org/kvm/36863cf7-61b9-4885-946d-1608179fcab4@xxxxxxxxxxxxxxxxx/T/#u
> > > ---
> > > arch/x86/include/asm/kvm_host.h | 7 +++
> > > arch/x86/include/asm/kvm_para.h | 19 ++++++++
> > > arch/x86/include/uapi/asm/kvm_para.h | 24 ++++++++++
> > > arch/x86/kernel/kvm.c | 68 +++++++++++++++++++++++++++
> > > arch/x86/kernel/process.c | 7 +++
> > > arch/x86/kvm/cpuid.c | 3 +-
> > > arch/x86/kvm/x86.c | 69 ++++++++++++++++++++++++++++
> > > 7 files changed, 196 insertions(+), 1 deletion(-)
> >
> > I really, really, reaaaaally don't want to take on any (more) PV scheduling ABI
> > in KVM, if possible. And I definitely don't want to take a bunch of one-off
> > hooks, e.g. for idle tracking and then for something else a few months/years
> > later.
> >
> > Vineeth and Josh are working on PV scheduling via sched_ext and I assume additional
> > communication channels. I'm guessing idle state is one of the things that'll get
> > communicated to the host? I don't know the exact status of their work, but they've
> > got a slot in the sched_ext MC at LPC:
> >
> > https://lpc.events/event/20/contributions/2482
>
> Thanks Sean, yea this looks achievable though the pvsched framework.
> Note that usage of sched_ext is not a requirement. For something like
> this idle time tracking, it would be trivial to extend the pvsched
> guest component, which already tracks guest scheduling state changes
> via guest context switch.
I also agree, this could be achieved with pvsched. In the latest
pvsched PoC (not yet posted upstream), the guest vcpu infact updates
pvsched shm when it enters idle. Host side of the pvsched(default
pvsched policy) uses this information to take some actions before the
vcpu thread sleeps in the host. Idle time accounting is also not yet
implemented, but it would be trivial to build on top of this
framework.
Thanks,
Vineeth