Re: [PATCH 07/11] PCI: rcar-gen4: Recover the Root Port on link down
From: Marek Vasut
Date: Sun Sep 27 2026 - 18:38:36 EST
On 9/24/26 6:15 PM, Koichiro Den wrote:
On Tue, Sep 22, 2026 at 11:44:02PM +0200, Marek Vasut wrote:
On 9/18/26 5:20 AM, Koichiro Den wrote:
On R-Car, intreq_pcim_sub carries both the integrated MSI receiver and
the controller's reset requests (smlh_req_rst_not, link_req_rst_not), so
the generic DesignWare chained handler reads the MSI status from DBI as
soon as the link goes down. On R-Car S4 that is a hazard: DBI accesses
issued within a few hundred microseconds of an unexpected link down do
not complete and hang the host. In testing, the first Root Port config
read after powering off the link partner hung unless delayed by ~300 us.
Out of curiosity, do they trigger SError, and does the firmware (TFA)
trap/fix those up in EL3?
I haven't confirmed whether those DBI accesses trigger an SError.
The BL31 on the Spider I'm using should be based on rcar-s4_v2.5 [1] from the
Spider BSP (as I haven't updated it). I just checked that this version also sets
HANDLE_EA_EL3_FIRST=1, plus plat_ea_handler() is empty. So IIUC an SError would
be taken to EL3
Yes.
, and the normal path would return to the kernel without fixing
the underlying error.
That is correct.
R-Car Gen3 does PCIe link error fixup in TFA, it looks this way:
https://github.com/ARM-software/arm-trusted-firmware/commit/0969397f295621aa26b3d14b76dd397d22be58bf
And, regardless of whether it's a synchronous abort or an SError:
- There was no crash report (via report_unhandled_exception) on the console.
- Once that DBI access in the small window hangs, it never recovers,
whereas adding udelay(300) before the access avoids the hang. So I
doubt it's just retrying the same access after a transient synchronous
abort as well (though I can't rule it out).
( It just crossed my mind, this sounds similar to 0056d29f8c1b ("PCI: rcar-gen4: Assure reset occurs before DBI access") )
[1] https://github.com/renesas-rcar/arm-trusted-firmware/tree/rcar-s4_v2.5
Use the pre-MSI callback to check the APP reset status before DBI is
touched. When a reset request is latched, mask the sources, ack the
request and schedule recovery work. The work calls
pci_host_handle_link_down(), which runs the AER-style recovery and
resets the controller through reset_root_port(). If the reset fails, the
sources stay masked so nothing touches the unrecovered controller.
Only unmasked status bits are handled and pending latches are cleared
when re-arming, so requests recorded during probe or the reset itself do
not trigger another recovery. Teardown only disables link-down
detection: MSI delivery has to keep working while devices are removed.
When iMSI-RX is not used (external MSI controller or pci=nomsi), the
DesignWare core does not request intreq_pcim_sub, so request it in the
driver.
Would it make sense to request the line unconditionally, to simplify the
driver(s) ?
Yes. The pcie-rcar-gen4 can set pp->msi_irq[0] = -ENODEV just like spear13xx,
dra7xx and keembay for the same reason. That also removes the need for adding
.pre_msi_irq addition in patch 4. Some other adjustments might be needed, but I
believe that is the right and clean direction. I will respin with that in mind.
Thanks for the review, that helps a lot!
Likewise, thank you for your help !
--
Best regards,
Marek Vasut