Re: [PATCH v6 0/4] KVM: PPC: Expose CPU compatibility modes for nested guests
From: Amit Machhiwal
Date: Thu Aug 06 2026 - 01:33:03 EST
Hi Ritesh,
Thanks for taking a look. Please find my response inline.
On 2026/08/06 12:09 AM, Ritesh Harjani wrote:
>
> Hi Amit,
>
> Amit Machhiwal <amachhiw@xxxxxxxxxxxxx> writes:
>
> > On POWER systems, newer processor generations can operate in compatibility
> > modes corresponding to earlier generations (e.g., a Power11 system running
> > in Power10 compatibility mode). In such cases, the effective CPU level
> > exposed to guests differs from the physical processor generation.
> >
> > This creates a problem for nested virtualization. When booting a nested KVM
> > guest (L2) inside a host KVM guest (L1) running in a compatibility mode,
> > userspace (e.g., QEMU) may derive the CPU model from the raw hardware PVR
> > and attempt to configure the nested guest accordingly. However, the L1
> > partition is constrained by the compatibility level negotiated with the
> > hypervisor (L0), and requests exceeding that level are rejected, leading to
> > guest boot failures such as:
> >
> > KVM-NESTEDv2: couldn't set guest wide elements
> >
> > This series provides a mechanism for userspace to query the effective CPU
> > compatibility modes supported by the host, so it can select an appropriate
> > CPU model for nested guests.
> >
> > To achieve this, the series introduces a new KVM capability and ioctl
> > (KVM_CAP_PPC_COMPAT_CAPS / KVM_PPC_GET_COMPAT_CAPS) that expose the
> > compatibility modes supported by the host.
> >
>
> Sorry, but I am somehow not convinced on whether we need all of this
> machinary just to get these 3 bits of information, which we are
> returning today.
>
> Since KVM_CHECK_EXTENSION can already return an int, so why can't we use
> KVM_CAP_PPC_COMPAT_CAPS itself and return the bitmap of supported compat
> modes to the user?
> Say if the cap is not supported, we can return 0, otherwise we can
> return the bitmap of supported compat modes. This will easily allow us
> to use 31-bits which as I see would be hardly a problem in the near
> future. In the future if it grows - we can always use KVM_CAP_PPC_COMPAT_CAPS2.
>
> This should reduce the code complexity both in the kernel and
> userspace and we don't even need a new ioctl then.
Thanks for the suggestion. I considered this approach but would like to
go with a dedicated ioctl for the following reasons:
1. Intended semantics: The KVM API documentation states:
..kvm defines extension identifiers and a facility to query
whether a particular extension identifier is available. If it is, a
set of ioctls is available for application use.
[...]
KVM defines many constants of the form KVM_CAP_*, each corresponding
to a set of functionality provided by one or more ioctls. Availability
of these capabilities can be checked with KVM_CHECK_EXTENSION.
The intended role of KVM_CAP_* is to signal ioctl availability, not
to serve as a data retrieval mechanism itself.
You may take a look at KVM_CAP_PPC_GET_CPU_CHAR for instance.
2. Return type constraint: KVM_CHECK_EXTENSION returns a signed 32-bit int. The
capability bits are defined as (1ULL << 62), (1ULL << 61), and (1ULL << 60) —
64-bit values that cannot fit in a 32-bit return. Renumbering them to small
integers would be a UAPI change and would lose alignment with the
H_GUEST_CAP_* values from the hypervisor ABI.
3. Extensibility: The struct-based approach with the size field provides clean
forward and backward ABI versioning via copy_struct_from/to_user(), without
needing a KVM_CAP_PPC_COMPAT_CAPS2 in the future.
Thanks,
Amit