Re: [PATCH v2] PCI: qcom: Block accesses to downstream devices on link down

From: Qiang Yu

Date: Thu Oct 01 2026 - 10:10:27 EST


On Wed, Sep 30, 2026 at 09:51:57AM -0700, Mayank Rana wrote:
> Hi Qiang
>
> On 9/13/2026 11:27 PM, Qiang Yu wrote:
> > On Fri, Sep 11, 2026 at 12:18:07PM -0500, Bjorn Helgaas wrote:
> > > On Fri, Sep 11, 2026 at 08:17:29AM +0200, Manivannan Sadhasivam wrote:
> > > > On Tue, Sep 08, 2026 at 06:00:30PM -0500, Bjorn Helgaas wrote:
> > > > > On Wed, Aug 19, 2026 at 11:36:54PM -0700, Qiang Yu wrote:
> > > > > > After a PCIe link goes down, software may still access the BAR
> > > > > > (MMIO) space or configuration space of devices behind that link
> > > > > > before recovery has run. As the link is down, these accesses
> > > > > > never complete, resulting in a storm of Completion Timeout AERs.
> > > > >
> > > > > What is special about qcom here? It seems like the Completion
> > > > > Timeouts and AER interrupts should happen with every PCIe
> > > > > controller.
> > > >
> > > > The special behavior which is common across many (not all) ARM SoCs
> > > > is that they don't synthesize all-one response for completion
> > > > timeouts, unlike RCs in x86 machines. Rather, they return AXI error
> > > > response, resulting in CPU treating them as SError, in-addition to
> > > > AER storm.
> > > >
> > > > Commit message missed mentioning SError though.
> > >
> > > I don't know how SError works, but this sounds like a pretty big open
> > > issue with respect to RAS. I don't think we want a kernel panic
> > > because a device failed to respond to a config or MMIO access, e.g.,
> > > if a card or Thunderbolt cable got unplugged.
> > >
> > > > > Is this mitigating an issue that will still happen on other
> > > > > controllers and should be solved elsewhere, e.g., by changing the
> > > > > software that accesses the BAR to look for the error responses it
> > > > > gets when the Completion Timeout happens?
> > > >
> > > > That would be too late as we don't get a proper error response.
> > > > That's why this patch is used.
> > > >
> > > > > The patch refers to the ECAM blocker (which I assume affects
> > > > > config accesses) and doesn't mention MMIO. Is the SLV_AXI stuff
> > > > > for MMIO?
> > > >
> > > > ECAM blocker is a Qcom's custom implementation which prevents the
> > > > config access to go out of the link and terminate the request
> > > > properly. And yes, it doesn't affect MMIO like BAR.
> > >
> > > The commit log implies that the patch blocks both config and MMIO
> > > accesses: "software may access MMIO or config space ... Use ECAM
> > > blocker to drop these accesses." If the ECAM blocker only affects
> > > config accesses, we need to reword that description.
> > >
> >
> > This patch can block both config space and MMIO accesses.
> To avoid confusion, can you rename APIs by replacing ecam by outbound
> i.e.
> qcom_pcie_init_ecam_blocker() as qcom_pcie_init_outboud_blocker() and
> qcom_pcie_enable_ecam_blocker() as qcom_pcie_enable_outbound_blocker().

Agreed, "outbound" is a better name here since this blocker drops both
config and MMIO accesses, not just config space as "ecam" implies.
This also lines up with Bjorn's earlier point that the commit log and
the API naming were inconsistent about scope.

Mani, this patch is applied to your branch, can I still respin this patch
to rename APIs?

- Qiang Yu