Re: [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup

From: Borislav Petkov

Date: Thu Oct 01 2026 - 20:01:19 EST


On Tue, Sep 29, 2026 at 01:25:09PM -0400, Zack Rusin wrote:
> Partly. So steal time tells Linux how much time a virtual CPU spent
> ready to run but waiting for the hypervisor to schedule it.
>
> For example, during a 100 ms interval, the vCPU might execute for 70
> ms and wait for a physical CPU for 30 ms. Reporting those 30 ms helps
> kernel:
> - account for CPU usage accurately: avoid charging applications for
> time when the hypervisor wasn't running their vCPU.
> - make fairer scheduling decisions: base task execution accounting on
> the CPU time tasks actually received.
> - expose host contention: the st field in top and counters in
> /proc/stat help explain why a VM is slow even though its applications
> don't appear to consume all available CPU time.

Aaaha, IOW, that's the "st" column here:

$ vmstat
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu-------
r b swpd free buff cache si so bi bo in cs us sy id wa st gu
1 0 0 1714852 9748 85300 0 0 3793 51 1445 0 0 1 99 0 0 0

In any case, you could keep that helpful explanation in yout 0th message. :)

> If options are binary then "no" :) It's not for monitoring tools, it's
> for the kernel's own steal-time accounting above, which currently
> doesn't work in confidential VMware guests. So the issue is that in a
> confidential guest (AMD SEV SNP, Intel TDX) memory is private by
> default, so the hypervisor can't read or write it. Any page the
> hypervisor has to write must first be converted to shared by the
> guest. Linux gives us the address of the steal-time buffer without
> converting it. The hypervisor then tries to write into a private page,
> which doesn't work. Currently ESXi deliberately powers off the VM when
> that happens. That only affects VMs with the steal clock enabled,
> which ESXi leaves off by default except for Photon guests. I'll
> probably change ESXi to just disable steal time when a guest gives us
> a private page, but that only avoids the power-off; steal time still
> won't work in these guests without this series.
>
> KVM has the same need for three of its per-CPU buffers (steal time,
> async page faults and PV EOI). It already converts them, but only on
> AMD and in its own loop. Kiryl asked on v1 for one common place that
> converts all such per CPU buffers early, on both AMD and Intel,

Right, why early?

I mean, I am still trying to see the justification for this diffstat

14 files changed, 247 insertions(+), 51 deletions(-)

and whether it is really worth it.

> instead of each hypervisor driver doing it. That's patches 2-6.
>
> Patch 1 fixes an old layout bug in uniprocessor kernels, where these
> buffers can share a page with unrelated data. The series also fixes
> SEV and SEV-SNP guests on KVM running kernels built with CONFIG_SMP=n,
> which currently hang at boot (we reproduced the hang).

That should tell you how much we care about UP. We would even love to make SMP
the default.

> Fair enough. I'll rework the cover letter and the commit messages so
> each one starts with the problem. The series originally wasn't really
> touching x86 core parts and I haven't updated it for a larger crowd.
> Would you like an explanation of steal time, like the above, in the
> cover as well?

Yes please.

Also, we have some blurb about how to write those:

https://docs.kernel.org/process/maintainer-tip.html#patch-subject

and

https://docs.kernel.org/process/submitting-patches.html

In talking to Peter about it, we were wondering whether this can be made
simpler. Like do not touch perCPU but do a normal page for each CPU's steal
time gunk and thus do not split the large page and then that early
enc/decrypting of memory I don't like either.

Perhaps we should start with the simplest approach first.

Thx.

--
Regards/Gruss,
Boris.

https://people.kernel.org/tglx/notes-about-netiquette