Re: [PATCH v20 06/14] KVM: arm64: Validate GCS exception lock when emulating ERET

From: Leonardo Bras

Date: Fri Sep 04 2026 - 09:38:57 EST


On Thu, Sep 03, 2026 at 08:22:44PM +0100, Mark Brown wrote:
> On Thu, Sep 03, 2026 at 04:37:37PM +0100, Leonardo Bras wrote:
> > On Tue, Sep 01, 2026 at 10:47:04PM +0100, Mark Brown wrote:
>
> > > + return vcpu_read_sys_reg(vcpu, GCSCR_EL2) & GCSCR_ELx_EXLOCKEN;
> > > +}
>
> > The above perfectly translates the GCS part of IllegalExceptionReturn().
>
> > It's a nit, as I suppose there should be no compiler warning on that, but
> > the function should return a bool, and the last return line returns an u64.
>
> > Maybe adding a "return !!()" would be better?
>
> There's no need to manually do translations like that in C, the
> conversion of 0 to false and any non-zero value to true when an
> integer is used in a boolean context is in the standard.
>

Yeah, I am aware the conversion will happen anyway, but I remember someone
complaining about something like this in the past, that's why I sent as a
nit.

> > > * - trying to return to EL1 with HCR_EL2.TGE set
> > > + * - GCSCR_ELx.EXLOCKEN is 1 and PSTATE.EXLOCK is 0 when attempting
> > > + * to return from ELx the same EL.
> > > */
> > > if (mode == PSR_MODE_EL3t || mode == PSR_MODE_EL3h ||
> > > mode == 0b00001 || (mode & BIT(1)) ||
> > > (spsr & PSR_MODE32_BIT) ||
> > > + kvm_check_illegal_exlock_return(vcpu, spsr) ||
> > > (vcpu_el2_tge_is_set(vcpu) && (mode == PSR_MODE_EL1t ||
> > > mode == PSR_MODE_EL1h))) {
> > > u64 mask;
> >
> > In IllegalExceptionReturn(), the GCS-related clause happens at the end, and
> > here it happens before the TGE one. Could this cause any weird behavior in
> > the future?
> >
> > I get that by doing like this you don't change the last line of the "if",
> > but I wonder if that could change anything.
>
> Given that we take the same action regardless of which or clause
> triggers I can't see how it would matter.

I see... well, as long as neither test ever have any collateral effect, I
think it should not matter, then.

>
> > > --- a/arch/arm64/kvm/hyp/vhe/switch.c
> > > +++ b/arch/arm64/kvm/hyp/vhe/switch.c
> > > @@ -383,6 +383,10 @@ static bool kvm_hyp_handle_eret(struct kvm_vcpu *vcpu, u64 *exit_code)
> > > return false;
> > > }
>
> > > + /* Push GCS exception lock failures into the slow path */
> > > + if (kvm_check_illegal_exlock_return(vcpu, spsr))
> > > + return false;
>
> > > /* If ERETAx fails, take the slow path */
> > > if (esr_iss_is_eretax(esr)) {
> > > if (!(vcpu_has_ptrauth(vcpu) && kvm_auth_eretax(vcpu, &elr)))
>
> > IIUC, this function will be called on the __kvm_vcpu_run_vhe() inner loop,
> > in the cases where the guest exited due to a eret.
>
> > What you change here is that in case of an illegal exlock return, it goes
> > out of the loop and return to host kernel, probably to deal with it in the
> > mentioned slowpath, the same way the ERETAx entry does.
>
> > I don't question on this being needed.
> > I would just like to understand why this is needed here.
>
> > This does not seem to be related to nested, as this is called in
> > __fixup_guest_exit() and not in fixup_nv_guest_exit(). But would not
> > hardware be responsible for cheking this, then?
>
> The code is here because it's part of the ERET handling, this should
> only happen for NV as we're not trapping ERET instructions otherwise.

Oh, makes sense!

> We need this because ERETs from vEL2 are handled in software, modulo the
> NV3 fast path mentioned at the top of the function.

Oh, and this is done in __fixup_guest_exit() because vEL2 is not a
nested guest. It would be it's guests' exit that would be dealt in
fixup_nv_guest_exit().

Is this correct?

Thanks!
Leo