Re: [PATCH] PCI: fix use-after-free in pci_pme_list_scan()
From: Christian Hedin
Date: Thu Oct 01 2026 - 03:31:47 EST
Hi Torsten,
I hit what looks like the same use-after-free on an unpatched, non-PaX
kernel, so here is a second data point in case it helps the patch along.
Hardware: Framework Laptop 13 (AMD Ryzen AI 300 Series), BIOS 03.05
Kernel: 7.2.5 (distro build, 7.2.5-3-omarchy), not tainted
Device: LG 40WT95UF monitor on USB4, with PCIe and USB tunneled
The monitor went into its automatic standby overnight (laptop awake, lid
closed), which drops the USB4 link and hot-removes the tunneled PCIe
bridges. In the same second as the removal:
thunderbolt 0-2: device disconnected
pcieport 0000:00:01.1: pciehp: Slot(0): Card not present
pci_bus 0000:03: busn_res: [bus 03-21] is released
pci_bus 0000:22: busn_res: [bus 22-40] is released
pci_bus 0000:41: busn_res: [bus 41-5f] is released
pci_bus 0000:02: busn_res: [bus 02-5f] is released
BUG: unable to handle page fault for address: 0000075700000060
#PF: supervisor read access in kernel mode
#PF: error_code(0x0000) - not-present page
Oops: Oops: 0000 [#1] SMP NOPTI
CPU: 10 UID: 0 PID: 4136500 Comm: kworker/10:1 Not tainted 7.2.5-3-omarchy #1 PREEMPT(full)
Workqueue: events_freezable pci_pme_list_scan
RIP: 0010:pci_pme_list_scan+0x4e/0x220
Code: ... 48 8b 43 10 <4c> 8b 60 38 4d 85 e4 0f 84 b8 00 00 00 ...
RAX: 0000075700000028 RBX: ffff8b26f0cca000
Call Trace:
process_one_work+0x19f/0x370
worker_thread+0x1b1/0x330
kthread+0xe4/0x120
ret_from_fork+0x2bd/0x350
ret_from_fork_asm+0x1a/0x30
note: kworker/10:1[4136500] exited with irqs disabled
It faults on the same load as your trace (bus->self, 0x38 off RAX), with a
garbage pdev->bus in RAX instead of the 0xfe poison, i.e. the pci_dev at RBX had already been freed and reused.
Without free poisoning the stale pointer is just whatever landed there,
which is probably why this is rarely seen.
One thing worth adding to the commit message: without panic_on_oops the
damage is not limited to the one worker. It dies while holding
pci_pme_list_mutex, so every later pci_pme_active() blocks forever. Here
that meant:
- irq/34-pciehp and several pm workqueue workers stuck in D state, so
the monitor's USB and PCIe functions never came back on replug (DP
tunnelling still worked).
- Any config space read that needs a runtime resume hangs unkillably:
task:lspci state:D
__mutex_lock.constprop.0+0x3e6/0x930
pci_pme_active+0x158/0x1f0
__pci_enable_wake+0x90/0xc0
pci_pm_runtime_resume+0x9a/0x130
...
pci_read_config+0x98/0x310
- System suspend failed in a loop for four hours ("Freezing user space
processes failed ... 3 tasks refusing to freeze"), and a clean reboot
hung as well.
So on a stock kernel a single hot-unplug can leave the machine needing a
hard power-off, which may be an argument for Cc: stable.
It is a race here: the same boot had two earlier disconnects of the same
device without an oops.
I can test a v2 on this machine if that is useful, though reproducing may
take a while given how rarely it fires. Full kernel log available on
request.
Thanks,
Christian