Re: [PATCH RFC 2/2] sched/core: Defer preempted remote vCPU task clock updates
From: Dongli Zhang
Date: Wed Sep 23 2026 - 13:32:30 EST
On Mon, Sep 21, 2026 8:59:08AM -0700, Sean Christopherson wrote:
> On Sun, Aug 23, 2026, Dongli Zhang wrote:
>> A remote update of a runqueue can advance rq->clock while the owner
>> vCPU is still preempted by the host. KVM publishes the matching stealtime
>> when the vCPU is about to re-enter the guest, so the remote CPU can
>> otherwise charge the stolen interval to rq->clock_task.
>>
>> Defer clock_task updates made by a remote CPU while the owner vCPU is
>> reported preempted. Fold the deferred delta back into the next update
>> that can proceed so IRQ and steal accounting process it together.
>>
>> This requires the hypervisor to publish up-to-date stealtime before
>> clearing the preempted data.
>
> What happens if the hypervisor doesn't do that? Because it's infeasible to
> guarantee this will never run on an older version of KVM.
>
With an older KVM, st->preempted can be cleared before the corresponding
stealtime update is published.
As a result, the guest can observe !vcpu_is_preempted() and consume the deferred
clock delta while still reading the old stealtime value. Thus, the stolen
interval is incorrectly charged to rq->clock_task in that update.
On the next update, the accumulated stealtime delta is observed, but it can be
larger than the new, small rq clock delta and is capped to that delta. Since
prev_steal_time_rq is nevertheless advanced to the new stealtime value, the
remaining stealtime is not accounted in later updates.
That's why I tagged this patchset as RFC, to seek feedback on whether there is a
better solution.
Thank you very much!
Dongli Zhang