[RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools

From: Luigi Rizzo

Date: Sat Oct 03 2026 - 17:22:53 EST


Here is a subsystem called DMA_PMD on which I would like feedback on
architecture, possible enhancements, or kernel components that could be
reused to avoid duplication.

I think it can be extremely useful for those affected by the HW/SW overhead
of IOMMU (especially IOTLB thrashing, see [1]), or confidential computing,
or preemptable VMs.

The (not too exaggerated) pitch line is

DMA_PMD has the performance of identity, guarantees that only IO buffers
can ever have active IOMMU or IOTLB mappings, removes the bounce buffer
overhead in confidential computing and preemptable VMs, and integrates
smoothly with existing kernel APIs.

=== ARCHITECTURE

DMA_PMD was initially designed to address IOTLB thrashing (details in [1]),
but turned out to also resolve nicely the strict IOMMU overhead, and avoid
the bounce buffer overhead in preemptible VMs and Confidential Computing.

It works by combining several known techniques:

- transparently feed allocators of device memory (skb_page_frag_refill(),
dma_alloc_attrs(), pagepool, ...) with DMA_PMD pages, i.e. physically
contiguous PMD_SIZE (2MB) pages mapped via PDE_SIZE IOMMU entries

- heavy recycling of DMA_PMD pages (like pagepool) and lazy IOMMU unmapping
(like DMA-FQ), BUT:

- like strict IOMMU (DMA), safely release memory back to the kernel only
after destroying all of its IOMMU mappings and a synchronous IOTLB flush

- on allocation, DMA_PMD pages can be configured to be pinned in the host
(hence suitable for preemptible VMs) and/or unencrypted (hence suitable
for Confidential Computing), removing the need for bounce buffers

- there are global and per-device optin /sys/device/.../dma_pmd_*
and /proc/sys/net/core/tx_enable_dma_pmd

NIC drivers can typically use dma_pmd with no modifications for rings,
tx buffers, rx buffers (if they use pagepool), and rx headers. For tx
headers, almost all drivers need changes (see later patches in the series)
to implement cheap bounce buffers and avoid individual 4K mappings.

The series has the following main components:

- a sparse array (similar to pageblock_flags) to quickly attach metadata
to a 2MB page without fiddling with the struct page. Cost is 128B per 2MB
page used as an IO buffer, totally negligible.

- dma_pmd_pool, a replacement for alloc_pages() that can be instantiated
per-CPU or per receive queue. It handles the split of PMD_SIZE pages
into order-N blocks, handles dma_map and unmap, and aggressively recycles
entries. It is used to feed skb_page_frag_refill(), pagepool, and
receive buffers for drivers that do not use pagepool.
See [2] for "WHY NOT PAGEPOOL FOR TX AND EVERYTHING"

- dma_pmd_arena, is another allocator backed by DMA_PMD pages and
is used exclusively as the backend for dma_alloc_attrs().
Used for longer-lived allocations (descriptor/completion rings,
rx and tx header buffers)

- glue code to hook DMA_PMD into pagepool [2], skb_page_frag_refill(),
dma_alloc_attrs(), dma_map/unmap...

- per-driver patches, where necessary (e.g. tx header buffers [3])

=== PERFORMANCE BENEFITS

Your mileage may vary. Enabling the IOMMU may have no throughput
impact, until it does when some system components (bus, IOTLB, CPU)
become overloaded. Aside from throughput reduction, one interesting
parameter is the effectiveness of the IOTLB. Here is a sample of SMMU
performance counters for a large ARM system with 2x200G NICs doing
bidirectional traffic:

=== DMA-FQ MODE (total throughput ~440Gbps)
25,757,488 smmuv3_pmcg_*/event=0x80/ IOTLB lookups
16,634,941 smmuv3_pmcg_*/event=0x81/ IOTLB misses

=== DMA_PMD on top of strict DMA (total throughput ~745Gbps)
22,693,218 smmuv3_pmcg_*/event=0x80/ IOTLB lookups
30,536 smmuv3_pmcg_*/event=0x81/ IOTLB misses
(not a mistake, also lookups went down despite the higher rate
because the tx side can use larger segments)

The 2MB mappings made IOTLB misses almost non existent, because
the working set is reduced by a factor of ~512.

Note, the code has more verbose comments than I would like.
Several of them are there to avoid Sashiko getting confused and
flagging false positives.

=== NOTES

[1] IOMMU IMPACT AND IOTLB THRASHING
This is documented in more detail in Documentation/core-api/dma-pmd.rst
but the compact version is below.

There are three main costs involved with using the IOMMU:
- CPU cost for dma map/unmap
- CPU cost and latency for synchronous IOTLB flush (required for strong security)
- IOTLB thrashing, shows up dramatically when the IO access pattern
exceeds the IOTLB size, and IOMMU page walks slow down bus activity
up to a point where we see over 30..60% throughput reduction just for
this reason. Some relevant references

https://lore.kernel.org/all/4b42f2eb-dc29-153e-ace9-5584ea2e5070@xxxxxxxxxx/
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10280580/

Specifically, NICs access 4 or more pages per packet (descriptor queue,
completion queue, rx or tx header, one or more buffers), with very little
locality especially for data buffers and tx headers. Each queue has 1-4K
entries, a fast NIC uses 16 or more queues per direction, so just the
data buffers cover 32..128K pages. Using 4KB mappings makes the IOTLB
ineffective, as shown by the performance counters shown earlier.

[2] WHY NOT PAGEPOOL FOR TX AND EVERYTHING
There is some overlap between pagepool and dma_pmd_pool in that both
implement a cache, but pagepool is missing some important features:

- pagepool does not handle splitting 2M pages in smaller chunks, or
the mappings. That would need to be implemented (and is what dma_pmd_pool does)

- pagepool's provided pages cannot be cleanly managed by
get_page()/put_page(), which is what all the consumers of
skb_page_frag_refill(). Fixing that would require dozen of changes
with high risk of missing some paths.

- even for the receive path, a provider for pagepool has more restrictions
than alloc_pages-supplied memory, so if I wanted to wrap dma_pmd_pool
as a provider I would have needed changes to the provider interface.

- pagepool requires a device at allocation time, which is not known when
skb_page_frag_refill() runs, hence the IOMMU mappings cannot be handled
at allocation time.

- page_pool_alloc_pages assumes a single NAPI consumer with BH disabled,
whereas skb_page_frag_refill() runs from a preemptive context.

- and last not least, we need to handle the dma_alloc_attrs() allocations
which don't match pagepool API

Of course all of the above could be modified, but the complexity
would be similar to that of implementing dma_pmd_pool, and with a
huge risk of breaking existing functionality or missing some path.

[3] DRIVER CHANGES

Most drivers need no change for rings, transmit buffers, or receive buffers
(if they use pagepool). Tx headers are generally mapped on the fly and
that is enough to trigger IOTLB thrashing, so most drivers need a small
change to implement very cheap tx bounce buffers backed by DMA_PMD pages.

We could avoid the bounce buffers using DMA_PMD for the skb->head region,
but that memory is contiguous to skb_shinfo so it would be exposed to
IO device access, which may be undesirable for security.

Luigi Rizzo (22):
iommu/dma: introduce CONFIG_DMA_PMD and metadata table
iommu/dma: add DMA_PMD pool lifecycle and page recycle hook
mm: Add split_page_compound()
iommu/dma: add DMA_PMD pool block allocation
iommu/dma: Global cap and shrinker for DMA_PMD pool memory
iommu/dma: reserve a per-domain IOVA window for DMA_PMD pages
iommu/dma: release DMA_PMD domain mappings on domain teardown
iommu/dma: use per-domain IOVA window to map DMA_PMD memory
iommu/dma: Add DMA_PMD arena allocator
driver core: Add per-device dma_pmd_* sysfs attributes
dma-mapping: Use DMA_PMD arena for dma_alloc_attrs()
net/core: Use per-CPU DMA_PMD pools for skb_page_frag_refill()
net/core: Use DMA_PMD for page_pool memory
iommu/dma: Support decrypted and pinned DMA_PMD pages
iommu/dma: Add background page scrubber for DMA_PMD pools
iommu/dma: Add per-NUMA-node PMD page reservoir
net/gve: Use DMA_PMD memory for RX buffers
net/gve: Use DMA_PMD memory for tx header bounce buffers
net/mlx5e: Use DMA_PMD memory for tx header bounce buffers
net/idpf: Use DMA_PMD memory for tx header bounce buffers
net/bnxt: Use DMA_PMD memory for tx header bounce buffers
iommu/dma: Add DMA_PMD statistics and debugfs

Documentation/core-api/dma-pmd.rst | 320 ++++
Documentation/core-api/index.rst | 1 +
drivers/base/core.c | 51 +
drivers/iommu/Kconfig | 33 +
drivers/iommu/Makefile | 3 +
drivers/iommu/dma-iommu.c | 72 +-
drivers/iommu/dma-iommu.h | 8 +
drivers/iommu/dma-pmd-arena.c | 412 +++++
drivers/iommu/dma-pmd-kunit.c | 225 +++
drivers/iommu/dma-pmd-map.c | 531 ++++++
drivers/iommu/dma-pmd-meta.c | 403 +++++
drivers/iommu/dma-pmd-pool.c | 1538 +++++++++++++++++
drivers/iommu/dma-pmd-priv.h | 385 +++++
drivers/iommu/iommu.c | 2 +-
drivers/net/ethernet/broadcom/bnxt/bnxt.c | 45 +-
drivers/net/ethernet/broadcom/bnxt/bnxt.h | 2 +
drivers/net/ethernet/google/gve/gve.h | 5 +
drivers/net/ethernet/google/gve/gve_main.c | 5 +
drivers/net/ethernet/google/gve/gve_rx.c | 61 +-
drivers/net/ethernet/google/gve/gve_tx.c | 38 +-
drivers/net/ethernet/google/gve/gve_tx_dqo.c | 45 +-
.../ethernet/intel/idpf/idpf_singleq_txrx.c | 6 +-
drivers/net/ethernet/intel/idpf/idpf_txrx.c | 21 +-
drivers/net/ethernet/intel/idpf/idpf_txrx.h | 27 +-
drivers/net/ethernet/mellanox/mlx5/core/en.h | 2 +
.../net/ethernet/mellanox/mlx5/core/en/txrx.h | 4 +-
.../net/ethernet/mellanox/mlx5/core/en_main.c | 14 +
.../net/ethernet/mellanox/mlx5/core/en_tx.c | 28 +-
include/linux/device.h | 5 +
include/linux/dma-pmd.h | 198 +++
include/linux/mm.h | 2 +
include/net/libeth/tx.h | 5 +-
include/net/page_pool/types.h | 3 +
include/net/sock.h | 3 +
kernel/dma/direct.c | 6 +-
kernel/dma/mapping.c | 39 +-
mm/Kconfig.debug | 11 +
mm/Makefile | 1 +
mm/page_alloc.c | 89 +
mm/split_page_compound_kunit.c | 113 ++
net/core/page_pool.c | 61 +-
net/core/sock.c | 86 +-
net/core/sysctl_net_core.c | 7 +
43 files changed, 4835 insertions(+), 81 deletions(-)
create mode 100644 Documentation/core-api/dma-pmd.rst
create mode 100644 drivers/iommu/dma-pmd-arena.c
create mode 100644 drivers/iommu/dma-pmd-kunit.c
create mode 100644 drivers/iommu/dma-pmd-map.c
create mode 100644 drivers/iommu/dma-pmd-meta.c
create mode 100644 drivers/iommu/dma-pmd-pool.c
create mode 100644 drivers/iommu/dma-pmd-priv.h
create mode 100644 include/linux/dma-pmd.h
create mode 100644 mm/split_page_compound_kunit.c

--
2.56.0.rc1.315.gc6ed9934b7-goog