Re: [PATCH] PCI/P2PDMA: Add Intel Haswell client host bridge to the whitelist

From: Bjorn Helgaas

Date: Mon Oct 05 2026 - 11:06:50 EST


On Mon, Oct 05, 2026 at 09:26:51PM +0800, Scott Lee wrote:
> P2P DMA between two devices that do not share an upstream bridge is only
> permitted when the traffic goes through a host bridge that is listed in
> pci_p2pdma_whitelist[] with REQ_SAME_HOST_BRIDGE.
>
> The Haswell client host bridge (8086:0c08, e.g. H97/H87/B85) is not in that
> list, and cpu_supports_p2pdma() returns false on Intel, so
> calc_map_type_and_dist() refuses P2P for any pair of devices behind
> different root ports of the same host bridge, even though that root complex
> handles the traffic correctly.
>
> Measured on an ASUS H97-PRO with a Xeon E3-1231 v3 (Haswell, 8086:0c08 at
> 00:00.0) and two GPUs on separate root ports of that host bridge
> (00:01.0 and 00:1c.4):
>
> - pci_p2pdma_distance() < 0 before the change; peer access granted after
> (amdgpu stops reporting "PCIe P2P access ... is not supported by the
> chipset" and hipDeviceCanAccessPeer() returns true in both directions)
> - 10 KB cross-device copy: 62-75 us vs 130-131 us host-staged (~2x)
> - 512 MB cross-device copy verified byte-wise against a known buffer:
> no corruption
> - tensor-parallel LLM inference (>25 GB model across both GPUs):
> +53% tokens/s
>
> REQ_SAME_HOST_BRIDGE is the correct flag here: the two devices share the
> upstream bridge (two root ports of one host bridge) rather than sitting
> behind a PCIe switch.
>
> This entry only permits P2P DMA to be used on this platform, it does not
> force it: the existing checks on ACS redirect, on the map type and on
> pci_p2pdma_distance() all still apply. Tested on one board only, so the
> usual caveat applies - broad testing on other Haswell boards would be
> needed before this can be considered generally safe.
>
> Signed-off-by: Scott Lee <dsix123@xxxxxxxxx>

Applied to pci/p2pdma for v7.4, thanks!

> ---
> # --- evidence from the machine this was tested on (stripped by git am) ---
> #
> # Hardware: ASUS H97-PRO (BIOS 2906), Xeon E3-1231 v3 (Haswell), 32 GB RAM,
> # 2x AMD Radeon RX 9060 XT (16 GB each), no PCIe switch
> #
> # lspci -nn:
> # 00:00.0 Host bridge: Intel 4th Gen Core Processor DRAM Controller [8086:0c08]
> # 00:01.0 PCI bridge: CPU PEG port; GPU0 is 0000:03:00.0 behind it
> # 00:1c.4 PCI bridge: PCH root port; GPU1 is 0000:09:00.0 behind it
> #
> # unpatched kernel (7.0.0-34-generic, distro stock):
> # $ sudo dmesg | grep 'not supported by the chipset'
> # amdgpu 0000:09:00.0: PCIe P2P access from peer device 0000:03:00.0 is not
> # supported by the chipset
> #
> # patched kernel (7.0.0-99-generic, same source tree + only this hunk):
> # $ sudo dmesg | grep 'not supported by the chipset' (no output)
> # $ python3 p2p_verify.py
> # can_access_peer: True / True
> # peer 512 MiB copy, byte-wise compared against a known buffer: OK
> # $ python3 p2p_latency.py (10 KiB, HIP events)
> # peer 0->1: 62.1 us peer 1->0: 74.5 us
> # host 0->1: 130.0 us host 1->0: 131.0 us
> # $ llama-bench, 27B model split across both GPUs, tensor vs layer split:
> # 25.05 +/- 0.91 tok/s vs 16.37 tok/s (+53%), greedy output identical
> #
> # Note the second GPU sits behind the PCH (Gen2 x4) - the setup still works
> # correctly despite the asymmetric and slow link, which is what makes it
> # interesting for the whitelist argument.
> #
> drivers/pci/p2pdma.c | 2 ++
> 1 file changed, 2 insertions(+)
>
> diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
> index 9334eb3..4f58564 100644
> --- a/drivers/pci/p2pdma.c
> +++ b/drivers/pci/p2pdma.c
> @@ -545,6 +545,8 @@ static const struct pci_p2pdma_whitelist_entry {
> /* Intel Xeon E7 v3/Xeon E5 v3/Core i7 */
> {PCI_VENDOR_ID_INTEL, 0x2f00, REQ_SAME_HOST_BRIDGE},
> {PCI_VENDOR_ID_INTEL, 0x2f01, REQ_SAME_HOST_BRIDGE},
> + /* Intel Haswell (client) */
> + {PCI_VENDOR_ID_INTEL, 0x0c08, REQ_SAME_HOST_BRIDGE},
> /* Intel Skylake-E */
> {PCI_VENDOR_ID_INTEL, 0x2030, 0},
> {PCI_VENDOR_ID_INTEL, 0x2031, 0},
> --
> 2.43.0
>