Re: [PATCH v3 12/18] KVM: arm64: Prevent host PC adjustments for protected vCPUs

From: Marc Zyngier

Date: Mon Sep 14 2026 - 10:19:58 EST


On Mon, 14 Sep 2026 12:33:32 +0100,
Fuad Tabba <fuad.tabba@xxxxxxxxx> wrote:
>
> __kvm_adjust_pc() lets the host advance a vCPU's PC or inject an
> exception, which for a protected vCPU would let the host redirect
> guest execution. Drop the request there: the entry handlers apply the
> host's PC_UPDATE_REQ on re-entry, where EL2 allows it.
>
> __kvm_adjust_pc() adjusts the vCPU a get/put pair returns, the one it
> was given outside pKVM. Both host callers hold the vCPU mutex, so a
> hyp vCPU loaded for the vCPU is loaded on the calling CPU. The request
> for a loaded protected vCPU is dropped. For a loaded non-protected
> vCPU, PKVM_HOST_STATE_DIRTY selects the copy, since adjusting the hyp
> vCPU while the host copy is authoritative loses the update at the next
> flush. Adjusting the hyp vCPU copies PC_UPDATE_REQ in and back out
> again: without the copy back, INCREMENT_PC outlives the adjustment and
> the next KVM_SET_VCPU_EVENTS trips WARN_ON(INCREMENT_PC) in
> kvm_pend_exception(). With no hyp vCPU loaded, as under
> KVM_SET_VCPU_EVENTS, the host copy is adjusted as before.
>
> Until the marshalling patch clears PC_UPDATE_REQ on the host copy at
> exit, a KVM_RUN that returns to userspace with INCREMENT_PC set leaves
> it on the host copy of a loaded protected vCPU, and a
> KVM_SET_VCPU_EVENTS before the next run then trips that WARN_ON.
>
> Suggested-by: Marc Zyngier <maz@xxxxxxxxxx>
> Signed-off-by: Fuad Tabba <fuad.tabba@xxxxxxxxx>
> ---
> arch/arm64/kvm/hyp/exception.c | 21 +++++++++-----
> arch/arm64/kvm/hyp/include/hyp/adjust_pc.h | 19 +++++++++++++
> arch/arm64/kvm/hyp/nvhe/hyp-main.c | 32 ++++++++++++++++++++++
> 3 files changed, 65 insertions(+), 7 deletions(-)
>
> diff --git a/arch/arm64/kvm/hyp/exception.c b/arch/arm64/kvm/hyp/exception.c
> index 6e60d890afa4a..7ea62e5304ae9 100644
> --- a/arch/arm64/kvm/hyp/exception.c
> +++ b/arch/arm64/kvm/hyp/exception.c
> @@ -353,12 +353,19 @@ static void kvm_inject_exception(struct kvm_vcpu *vcpu)
> */
> void __kvm_adjust_pc(struct kvm_vcpu *vcpu)
> {
> - if (vcpu_get_flag(vcpu, PENDING_EXCEPTION)) {
> - kvm_inject_exception(vcpu);
> - vcpu_clear_flag(vcpu, PENDING_EXCEPTION);
> - vcpu_clear_flag(vcpu, EXCEPT_MASK);
> - } else if (vcpu_get_flag(vcpu, INCREMENT_PC)) {
> - kvm_skip_instr(vcpu);
> - vcpu_clear_flag(vcpu, INCREMENT_PC);
> + struct kvm_vcpu *target = pkvm_adjust_pc_get(vcpu);

nit: it is really odd to see this 'pkvm' prefix in generic code. The
point of it is not only to abstract the pvkm complexity away, but also
to have a wrapper that may be of use in other situations.

Don't respin the series just for this though, I may end-up changing
this when applying it.

> + if (!target)
> + return;
> +

I find this one scary, see below.

[...]

> diff --git a/arch/arm64/kvm/hyp/nvhe/hyp-main.c b/arch/arm64/kvm/hyp/nvhe/hyp-main.c
> index da8ab636063cf..1dcc75261dc08 100644
> --- a/arch/arm64/kvm/hyp/nvhe/hyp-main.c
> +++ b/arch/arm64/kvm/hyp/nvhe/hyp-main.c
> @@ -642,6 +642,38 @@ static void handle___pkvm_host_mkyoung_guest(struct kvm_cpu_context *host_ctxt)
> cpu_reg(host_ctxt, 1) = ret;
> }
>
> +/*
> + * PKVM_HOST_STATE_DIRTY names the authoritative copy, the host's when set.
> + * A loaded protected vCPU takes the request at its next entry instead.
> + */
> +struct kvm_vcpu *pkvm_adjust_pc_get(struct kvm_vcpu *vcpu)
> +{
> + struct pkvm_hyp_vcpu *hyp_vcpu;
> +
> + if (!is_protected_kvm_enabled())
> + return vcpu;
> +
> + hyp_vcpu = pkvm_get_loaded_hyp_vcpu();
> + if (!hyp_vcpu || hyp_vcpu->host_vcpu != vcpu)

Under which circumstances do we get hyp_vcpu->host_vcpu != vcpu?

> + return vcpu;
> +
> + if (pkvm_hyp_vcpu_is_protected(hyp_vcpu))
> + return NULL;

Is it always the case that a protected vcpu cannot see its PC adjusted
at all? How is PC updated after an exit for MMIO? I feel there is an
interaction with the above, but I'm not 100% certain...

> +
> + if (vcpu_get_flag(vcpu, PKVM_HOST_STATE_DIRTY))
> + return vcpu;
> +
> + vcpu_copy_flag(&hyp_vcpu->vcpu, vcpu, PC_UPDATE_REQ);
> + return &hyp_vcpu->vcpu;
> +}
> +
> +/* Reflect the consumed request back, otherwise it stays pending. */
> +void pkvm_adjust_pc_put(struct kvm_vcpu *vcpu, struct kvm_vcpu *target)
> +{
> + if (target != vcpu)
> + vcpu_copy_flag(vcpu, target, PC_UPDATE_REQ);
> +}
> +
> static void handle___kvm_adjust_pc(struct kvm_cpu_context *host_ctxt)
> {
> DECLARE_REG(struct kvm_vcpu *, vcpu, host_ctxt, 1);

Thanks,

M.

--
Without deviation from the norm, progress is not possible.