Re: [PATCH v2 00/15] PCI: rcar-gen4: Recover from link down and route Root Port interrupts

From: Marek Vasut

Date: Sat Oct 03 2026 - 14:30:35 EST


On 9/28/26 6:52 PM, Koichiro Den wrote:
Hi,

This series improves error handling on the R-Car Gen4 PCIe host
controller (tested on R-Car S4 Spider, r8a779f0). It fixes unexpected
link-down handling so the host doesn't hang, and wires up missing Root
Port interrupts (AER, PME, bandwidth notifications) so port services
actually work.

A few hardware quirks made this tricky:

1. PCIEINTSTS0 link-up bits don't track the actual link state. The
driver never noticed when the link dropped and kept trying config
accesses on the dead link. Patch 2 combines the APP link-up event
check from Figure 104.5 with the PORT_DEBUG1 live link check. Startup
clears the APP latches before enabling LTSSM, and the DWC core polls
the combined condition on the RC side.

2. On S4, a DBI access immediately after an unexpected link down can
hang the host. Commit 0056d29f8c1b ("PCI: rcar-gen4: Assure reset
occurs before DBI access") describes an SError on V4H after reset
deassertion, but whether the S4 hang shares the same underlying cause
has not been verified.

Adding a delay before the DBI access avoided the hang in my tests,
but that alone would not provide link-down recovery. Also, the reset
request shares intreq_pcim_sub with iMSI-RX, while AER arrives on
another IRQ. Delaying AER dispatch alone would therefore leave the
MSI handler exposed, and both paths would still need coordination
with the controller reset.

The driver handles the reset request itself and schedules
pci_host_handle_link_down(), as the rockchip and qcom drivers do.
This also provides recovery with older DTs that have no "aer"
interrupt, without relying on AER to initiate it. The reset callback
shares the reset sequence used at probe, including the reset-status
readback and delay added by 0056d29f8c1b.

3. Root Port interrupts only trigger APP status bits on platform IRQs.
The controller appears to lack SII2MSI, so Root Port MSIs never reach
the GIC ITS. The Root Port's INTx is routed to intreq_pcim_sub as
well, which the driver holds, so port services can't request it
either. This series works around it by hiding Root Port MSI caps
across the board and emulating INTx using a virtual IRQ domain, fed
by the "aer" IRQ and intreq_pcim_sub.

Patch 1 is only loosely related: it adds Renesas to the RAS DES VSEC
list so the DWC debugfs error injection works on R-Car. I used it to
test the Root Port AER path (see below) and included it here for that
reason. Happy to send it separately if preferred.

Based on next-20260925. The driver patches build on 7fc9907223d6 ("PCI:
rcar-gen4: Add Application/Local register reset control") in pci/next,
and the DTS patch is for renesas-devel, which already has b43aa6a6ebe8
("arm64: dts: renesas: r8a779f0: Add GICv3 ITS and update PCIe nodes");
configurations b and c below use that DT.

Note: backward compatibility with older DTs is kept. Without the "aer"
interrupt, only Root Port AER remains unavailable. See the Testing
section below.

Retesting with v2
-----------------

Setup: R-Car S4 Spider (RC) linked to another S4 Spider running the
pci-epf-test endpoint. pci_endpoint_test is bound on the RC side.

1. Link down / recovery. On the EP side, toggle the endpoint controller
off and on. The short pause keeps the endpoint away long enough for
the RC to notice, but brings it back within the reset window so
recovery can succeed. Adopted the test approach from [1]:

# cd /sys/kernel/config/pci_ep
# echo 0 > controllers/e65d0000.pcie-ep/start
# sleep 0.1
# echo 1 > controllers/e65d0000.pcie-ep/start

Expected on the RC dmesg:
pcieport 0000:00:00.0: Recovering Root Port due to Link Down
pcieport 0000:00:00.0: Root Port has been reset
pcieport 0000:00:00.0: AER: device recovery successful

and the "msi" (intreq_pcim_sub) interrupt count going up in
/proc/interrupts. Without this series nothing shows up here: the
link comes back on its own once the endpoint returns, but the RC
never notices the outage and the endpoint is left unconfigured
(see 4). Config accesses issued while the link is down hang the
host.

[1] https://lore.kernel.org/r/abFMa6DCGGLUHddA@fedora/

2. Bandwidth notification. On the RC, retrain the link:

# setpci -s 00:00.0 CAP_EXP+0x10.w=0x0c23

Expected:
- the virtual Root Port IRQ (rcar-gen4-rp in /proc/interrupts,
shared by PCIe PME, aerdrv and PCIe bwctrl) fires once
- bwctrl clears LnkSta.LBMS
(setpci -s 00:00.0 CAP_EXP+0x12.w reads 0x2024 again).

Before the series LnkSta read 0xe024 afterwards, LBMS and LABS
stuck.

3. Root Port AER. On the RC, inject an LCRC error with the DWC debugfs
(patch 1) and issue one config read so a TLP actually goes out:

# cd /sys/kernel/debug/dwc_pcie_e65d0000.pcie/rasdes_err_inj
# echo 1 > rx_lcrc # error detected by the Root Port
# setpci -s 01:00.0 VENDOR_ID.w
# echo 1 > tx_lcrc # error detected by the endpoint,
# setpci -s 01:00.0 VENDOR_ID.w # reported back with ERR_COR

Expected on the RC dmesg, respectively:
pcieport 0000:00:00.0: PCIe Bus Error: severity=Correctable
pcieport 0000:00:00.0: [ 6] BadTLP | Receiver | Data Link Layer

pcieport 0000:00:00.0: AER: Correctable Error message received from 0000:01:00.0
pci-endpoint-test 0000:01:00.0: PCIe Bus Error: severity=Correctable
pci-endpoint-test 0000:01:00.0: [ 6] BadTLP | Receiver | Data Link Layer

plus the virtual Root Port IRQ count and aer_rootport_total_err_cor
going up by one each time. The link stays up throughout, the DLL
retry recovers the TLP. Before the series nothing is reported.

4. Regression check. Run pci_endpoint_test after step 1. PASS/FAIL/SKIP
counts match a run without step 1.

Configurations:

a. Without this series**
b. GIC ITS, DT with the new "aer" interrupt (this series)
c. GIC ITS, DT without "aer" (b43aa6a6ebe8 ("arm64: dts: renesas:
r8a779f0: Add GICv3 ITS and update PCIe nodes") or later)
d. iMSI-RX, DT before b43aa6a6ebe8 (no msi-parent, no "aer")

Result:
recovery bwctrl/PME RP AER pcitest
---------------------------------------------------------
a. none no no all FAIL
b. ok ok ok no change
c. ok ok n/a* no change
d. ok ok n/a* no change

* Root Port AER needs the "aer" interrupt; without it the behaviour is
unchanged from before the series.
** Only patch 1 applied on top of the base, so the same debugfs error
injection could be used for the comparison. pci_endpoint_test fails
across the board there because nothing restores the endpoint after
the toggle.
This is one awesome cover letter.