Re: [PATCH v6 18/18] RDMA/mlx5: Ask P2PDMA whether ATS takes a direct peer-to-peer route
From: Thomas Hellström
Date: Thu Sep 17 2026 - 10:02:40 EST
Hi
On Mon, 2026-09-14 at 14:22 +0300, Leon Romanovsky wrote:
> From: Leon Romanovsky <leonro@xxxxxxxxxx>
>
> mlx5_umem_needs_ats() enables ATS for any dma-buf whose caller asked
> for
> Relaxed Ordering, on the assumption that a switch in the path has CR,
> RR
> and DT all set. It also enables it for a buffer already mapped with
> the
> peer's bus addresses, which are not translatable at all.
>
> P2PDMA has read the ACS controls, so ask it through
> dma_buf_p2pdma_map_type(): enable ATS only where the path is not
> routed
> directly as it stands, but would be for a Translated Request whose
> Completions carry Relaxed Ordering. Exporters that name no provider
> keep
> the old assumption, since their ACS settings remain hidden.
>
> Signed-off-by: Leon Romanovsky <leonro@xxxxxxxxxx>
> ---
> drivers/infiniband/hw/mlx5/mlx5_ib.h | 36 ++------------------------
> ------
> drivers/infiniband/hw/mlx5/mr.c | 40
> ++++++++++++++++++++++++++++++++++++
> 2 files changed, 42 insertions(+), 34 deletions(-)
>
> diff --git a/drivers/infiniband/hw/mlx5/mlx5_ib.h
> b/drivers/infiniband/hw/mlx5/mlx5_ib.h
> index e9ddf2e97a76..ab32742b2180 100644
> --- a/drivers/infiniband/hw/mlx5/mlx5_ib.h
> +++ b/drivers/infiniband/hw/mlx5/mlx5_ib.h
> @@ -1646,40 +1646,8 @@ static inline bool rt_supported(int ts_cap)
> ts_cap ==
> MLX5_TIMESTAMP_FORMAT_CAP_FREE_RUNNING_AND_REAL_TIME;
> }
>
> -/*
> - * PCI Peer to Peer is a trainwreck. If no switch is present then
> things
> - * sometimes work, depending on the pci_distance_p2p logic for
> excluding broken
> - * root complexes. However if a switch is present in the path, then
> things get
> - * really ugly depending on how the switch is setup. This table
> assumes that the
> - * root complex is strict and is validating that all req/reps are
> matches
> - * perfectly - so any scenario where it sees only half the
> transaction is a
> - * failure.
> - *
> - * CR/RR/DT ATS RO P2P
> - * 00X X X OK
> - * 010 X X fails (request is routed to root but root never
> sees comp)
> - * 011 0 X fails (request is routed to root but root never
> sees comp)
> - * 011 1 X OK
> - * 10X X 1 OK
> - * 101 X 0 fails (completion is routed to root but root
> didn't see req)
> - * 110 X 0 SLOW
> - * 111 0 0 SLOW
> - * 111 1 0 fails (completion is routed to root but root
> didn't see req)
> - * 111 1 1 OK
> - *
> - * Unfortunately we cannot reliably know if a switch is present or
> what the
> - * CR/RR/DT ACS settings are, as in a VM that is all hidden. Assume
> that
> - * CR/RR/DT is 111 if the ATS cap is enabled and follow the last
> three rows.
> - *
> - * For now assume if the umem is a dma_buf then it is P2P.
> - */
> -static inline bool mlx5_umem_needs_ats(struct mlx5_ib_dev *dev,
> - struct ib_umem *umem, int
> access_flags)
> -{
> - if (!MLX5_CAP_GEN(dev->mdev, ats) || !umem->is_dmabuf)
> - return false;
> - return access_flags & IB_ACCESS_RELAXED_ORDERING;
> -}
> +bool mlx5_umem_needs_ats(struct mlx5_ib_dev *dev, struct ib_umem
> *umem,
> + int access_flags);
>
> int set_roce_addr(struct mlx5_ib_dev *dev, u32 port_num,
> unsigned int index, const union ib_gid *gid,
> diff --git a/drivers/infiniband/hw/mlx5/mr.c
> b/drivers/infiniband/hw/mlx5/mr.c
> index 00e13028762a..286f372e5b0c 100644
> --- a/drivers/infiniband/hw/mlx5/mr.c
> +++ b/drivers/infiniband/hw/mlx5/mr.c
> @@ -38,6 +38,7 @@
> #include <linux/export.h>
> #include <linux/delay.h>
> #include <linux/dma-buf.h>
> +#include <linux/dma-buf-mapping.h>
> #include <linux/dma-resv.h>
> #include <rdma/frmr_pools.h>
> #include <rdma/ib_umem_odp.h>
> @@ -47,6 +48,45 @@
> #include "data_direct.h"
> #include "dmah.h"
>
> +MODULE_IMPORT_NS("DMA_BUF");
> +
> +bool mlx5_umem_needs_ats(struct mlx5_ib_dev *dev, struct ib_umem
> *umem,
> + int access_flags)
> +{
> + struct dma_buf_attachment *attach;
> +
> + if (!MLX5_CAP_GEN(dev->mdev, ats) || !umem->is_dmabuf)
> + return false;
> +
> + /*
> + * The Completer decides whether its Completions carry
> Relaxed
> + * Ordering, and only a Request that asked for it can expect
> them to.
> + */
> + if (!(access_flags & IB_ACCESS_RELAXED_ORDERING))
> + return false;
> +
> + attach = to_ib_umem_dmabuf(umem)->attach;
> + switch (dma_buf_p2pdma_map_type(attach, 0)) {
> + case PCI_P2PDMA_MAP_NONE:
> + /* Nothing is known about the route, so fall back to
> the bet. */
> + return true;
> + case PCI_P2PDMA_MAP_BUS_ADDR:
> + /*
> + * The path is routed directly already and is
> programmed with
> + * the peer's bus addresses. Those are not
> translatable, so
> + * ATS would be wrong as well as pointless.
> + */
> + return false;
> + default:
> + break;
> + }
> +
> + return dma_buf_p2pdma_map_type(attach,
> + PCI_P2PDMA_TLP_TRANSLATED |
> +
> PCI_P2PDMA_TLP_RELAXED_CPL) ==
> + PCI_P2PDMA_MAP_BUS_ADDR;
> +}
> +
It looks like this works well when the device can choose whether to
enable ATS per transaction.
However, at least for the Intel GPUs, ATS enablement is based on the
PCIe-side enable bit.
This means that if p2pdma tells the dma-mapping layer to give an Xe
device a bus address rather than an IOVA, things break, while
pci_p2pdma_distance says everything is OK.
It looks like the infrastructure and solution added in this series is
targeted at fixing this on the device side by conditionally enabling
ATS. However I think we need to look also at having the computed
routing assume untranslated transactions using IOVA rather than bus
address.
That is, a flag to tell the topology check that some transactions
*will* take the host-bridge path due to IOVA being used, and that the
computations including pci_p2pdma_distance() need to check whether that
is possible (checking whitelist etc.) and return the corresponding
THRU_HOST_BRIDGE mapping type. Translated transactions taking a short-
cut using the bus-address would then be hidden from the driver.
Whether that is best done as a parameter to these functions or perhaps
as a flag in the PCI device, I'm not sure.
Thanks,
Thomas
> static int mkey_max_umr_order(struct mlx5_ib_dev *dev)
> {
> if (MLX5_CAP_GEN(dev->mdev,
> umr_extended_translation_offset))