Re: [PATCH v3] PCI: pciehp: Fix hotplug on Catlow Lake with unreliable PME status
From: Kuppuswamy Sathyanarayanan
Date: Wed Oct 07 2026 - 15:58:35 EST
Hi Lukas,
On 9/25/2026 11:35 AM, Lukas Wunner wrote:
> On Fri, Sep 25, 2026 at 10:38:08AM -0700, Kuppuswamy Sathyanarayanan wrote:
>> Once the port is in D3hot, pciehp has cleared HPIE and depends on PME.
>> That is where Catlow breaks. The PME interrupt arrives, but PME Status in
>> Root Status is never set. pcie_pme_irq() returns IRQ_NONE, the port stays
>> in D3hot and the hot-add event is lost.
>
> If a device below the Root Port (instead of the Root Port itself)
> signals PME, does the Root Port misbehave in the same way?
> I.e. is the PME Status bit clear in that case as well?
No, a downstream PME is reported correctly. I tested with a NIC at
0a:00.0 (8086:1533) connected below Root Port 00:1c.6 (8086:7a3e). The NIC
and the Port were both runtime suspended to D3hot, and the NIC woke on link
change. All six PMEs I captured had PME Status set and Requester ID 0x0a00,
which matches the NIC's BDF, and went through the normal path in
pcie_pme_handle_request().
>
> If so, the proper solution might be to add a quirk to the PME driver,
> not the PCIe hotplug driver.
Given the above, the problem is narrower than broken PME in general. It
only affects the PME the Port generates for its own hotplug event, and
hotplug is the only user of that. So both drivers are possible places for
the fix, and I would like your and Bjorn's preference before sending v5.
1. pciehp route (v4)
https://lore.kernel.org/linux-pci/20260323223056.3119060-1-sathyanarayanan.kuppuswamy@xxxxxxxxxxxxxxx/
A quirk sets PCI_DEV_FLAGS_PME_UNRELIABLE on the affected Ports, and
pciehp_disable_interrupt() skips clearing HPIE for them. The Port
still goes to D3hot, but hotplug events arrive as ordinary hotplug
interrupts and PME is not needed. The downside is that it changes the
suspend behaviour added by eb34da60edee, which was Bjorn's concern.
2. PME route
Keep the same quirk flag, leave pciehp alone, and handle it in
pcie_pme_irq(). If a quirked Port interrupts while not in D0 and PME
Status is clear, resume the Port. pciehp_runtime_resume() then
re-enables HPIE and pciehp_check_presence() finds the new card.
Downstream PMEs still set PME Status and take the normal path, and
pcie_pme_work_fn() is unchanged.
if (PCI_POSSIBLE_ERROR(rtsta)) {
spin_unlock_irqrestore(&data->lock, flags);
return IRQ_NONE;
}
if (!(rtsta & PCI_EXP_RTSTA_PME)) {
spin_unlock_irqrestore(&data->lock, flags);
/*
* Some Root Ports don't set PME Status for a PME they
* generate for their own hotplug events. While the Port
* is not in D0, PME is the only interrupt it can signal
* for such events, so resume it and let pciehp pick up
* the event.
*/
if ((port->dev_flags & PCI_DEV_FLAGS_PME_UNRELIABLE) &&
port->current_state != PCI_D0) {
pci_wakeup_event(port);
pm_request_resume(&port->dev);
return IRQ_HANDLED;
}
return IRQ_NONE;
}
>
> pcie_pme_irq() checks PME Status and bails out if it's not set.
> That would need an amendment such that Root Ports with broken PME
> would always assume it's set if they receive a PME. I think that
> would be safe because even though PME is shared with other interrupts
> such as hotplug, I think it's the only interrupt source once the port
> is in D3hot.
>
> There's another check for PME Status in pcie_pme_work_fn().
> This one is tricky because it uses the PME Status bit to jump
> out of the for-loop. Does the Root Port at least set the
> Requester ID to an appropriate value? If so maybe that can be
> used as an indicator whether the loop should be terminated.
>
> Or maybe the PME Pending bit can be used in lieu of PME Status?
>
In the failing case Requester ID stays 0000 and PME Pending is not set
either, so neither can be used in place of PME Status. That is why option
2 resumes the Port directly from pcie_pme_irq() and does not go through
pcie_pme_work_fn() at all.
> Thanks,
>
> Lukas
--
Sathyanarayanan Kuppuswamy
Linux Kernel Developer