Re: [PATCH 4/4] irqchip/gicv3: Add workaround for FUJITSU-MONAKA erratum E#030003
From: Kohei Enju
Date: Wed Oct 07 2026 - 04:19:00 EST
Hi Marc,
On 10/06 10:42, Marc Zyngier wrote:
> On Tue, 06 Oct 2026 10:08:19 +0100,
> Kohei Enju <enju.kohei@xxxxxxxxxxx> wrote:
> >
> > On 10/02 13:26, Marc Zyngier wrote:
> > > On Fri, 02 Oct 2026 11:26:52 +0100,
> > > Tomohiro Misono <misono.tomohiro@xxxxxxxxxxx> wrote:
> > > >
> > > > From: Kohei Enju <enju.kohei@xxxxxxxxxxx>
> > > >
> > > > On affected FUJITSU-MONAKA CPUs, an SGI generated by writing to
> > > > ICC_SGI0R_EL1, ICC_SGI1R_EL1, or ICC_ASGI1R_EL1 may be lost if the
> > > > operation races with CPU interface processing triggered by the arrival
> > > > of a higher-priority interrupt, an update to a pending interrupt, or a
> > > > transition of the PE to the Sleep state. When this occurs, the system
> > > > register write does not complete, causing the issuing core to hang.
> > > >
> > >
> > > Is the SGI lost? Or is the sender core hanging?
> >
> > Both: the SGI is lost, and the sysreg write does not complete, causing
> > the sending core to hang.
>
> I don't think the two are distinguishable. You might want to simplify
> the commit message to simply say that the CPU hangs.
>
> [...]
>
> > > > + */
> > > > + for_each_cpu(cpu, mask)
> > > > + gic_send_sgi_via_rdist(cpu, d->hwirq);
> > > > +
> > > > + /* Force the above writes to GICR_ISPENDR0 to be executed */
> > > > + dsb(st);
> > >
> > > This doesn't force things to be executed. This is about completion of
> > > the access, and with an nGnRE mapping, it doesn't enforce that the
> > > stores actually reach the RDs, only an arbitrary point in the memory
> > > subsystem. The only way to guarantee this is to perform a read-back.
> >
> > Understood.
> >
> > I hadn't considered this when writing v1, but on further reflection, I
> > don't think gic_ipi_send_mask() needs to wait for the target CPUs to
> > handle the IPIs.
>
> It's not about the target CPU handling the IPI, that'd be crazy. It is
> about making sure that the IPI request has been observed by the HW and
> that it is not going to race with something else.
>
> > Given that the GICv2 driver does not wait for MMIO
> > write completion either, I don't think we need to ensure completion of
> > these writes before returning.
>
> GICv2 has different architectural requirements (i.e. none). GICv3 is a
> bit clearer, see the requirements at the end of 12.1.3 in IHI0069H.b.
>
> >
> > If my understanding is correct, I'll remove the dsb(st) in v2.
> >
>
> As indicated in the spec, you either need nGnRnE+DSB, or a read-back.
> Given that Linux uses nGnRE for all device mappings, read-back is the
> only option here.
I understand that section 12.1.3 describes how software can guarantee
completion of a memory-mapped GIC write. With an nGnRE mapping, if the
GICR_ISPENDR0 writes must complete before gic_ipi_send_mask() returns, a
read-back from each target Redistributor is needed.
What I am trying to understand is whether completion at this point is an
architectural requirement for this particular write, or a requirement of
Linux's ipi_send_mask() interface.
I could not find an explicit completion requirement for ipi_send_mask()
itself. The normal GICv3 path also appears to return after an ISB,
without the DSB described in section 12.1.6 as necessary to guarantee
that the associated Redistributor has observed the write. This is why I
am unsure what completion guarantee is required here.
More specifically, I would like to understand:
- What could an outstanding IPI request race with?
- Which component must have observed the request before
gic_ipi_send_mask() returns: the CPU interface, the sending PE's
Redistributor, or the target Redistributor?
A pointer to the relevant architectural requirement or an example of
such a race in Linux would help me understand why completion is needed
here.
Thanks,
Kohei
>
> Thanks,
>
> M.
>
> --
> Jazz isn't dead. It just smells funny.