[PATCH 0/6] iommu/virtio: batch the mapping replay

From: Anlai Lu

Date: Sun Oct 04 2026 - 07:56:06 EST


The virtio-iommu device frees a domain together with its last endpoint, so
the first endpoint attaching to a domain that already has mappings must
rebuild it: viommu_replay_mappings() re-sends every mapping the driver
holds. That walk sent one synchronous request per mapping with
mappings_lock held - and with it interrupts disabled. The IRQ-off window
therefore grew with the number of mappings: 233 ms for 8192 4K mappings,
about 29 s for a million, long enough to add scheduling latency to
everything else running on the CPU.

The same code had two long-standing defects, fixed by the first patches:

- nr_endpoints was bumped only after the replay, without the lock, while
viommu_map_pages() read the same counter, also without the lock, to
decide whether to send a MAP itself or leave it to the replay. A map
racing the first attach could be sent by nobody and never reach the
device.

- viommu_unmap_pages() removed whole mappings from its tree but queued a
single UNMAP built from the range the caller asked for, after dropping
the lock: the device could keep translating memory the driver had
released, and a concurrent map could be sent ahead of the UNMAP. That
queueing also allocated on the unmap path, which iommu_unmap_nofail()
does not allow.

The series:

- 1/6 publishes the endpoint before the replay, tracks an unfinished
replay with replay_pending, and puts the endpoint back on the old
domain when the ATTACH or the replay fails - with a replay pending for
it if the device dropped its mappings - so a failed attach cannot leave
the endpoint count short, the endpoint unattached (bypassed), or the
endpoint recorded on a domain the core does not reference;
- 2/6 queues one UNMAP per contiguous run of removed mappings, in the
section that removes them, covering the mappings' own ranges;
- 3/6 allocates the UNMAP request together with the mapping, on the map
path, so the unmap path never allocates;
- 4/6 stops queueing and draining once the device is removed;
- 5/6 reads the endpoint count under the lock in the iotlb paths;
- 6/6 queues the replayed mappings, one per lock section, and waits once.

Measured with QEMU/KVM and iommufd, N=8192 4K mappings: max IRQ-off window
233.7 ms -> 1.4-2.5 ms, flat from N=64 to 32768; attach ioctl 218.9/249.0
-> 50.3/48.8 ms (48-67 ms across runs); viommu_send_req_sync() calls 8193
-> 1; the device receives exactly the same MAPs. With large pages the
same 256 MB attach is about 4 ms (197 device requests); the remaining cost
is the device processing one MAP per mapping, which the protocol
requires. The concurrent map-vs-first-attach test loses mappings on the
unpatched driver in several runs and none with the series; lockdep and
KCSAN report nothing. Allocation-failure and device-rejection paths were
exercised with test-only fault injection in QEMU.

Patches 1-4 carry Fixes: tags for the original driver commit. Behaviour
that predates this work - in particular, how a MAP the device rejects on
the map path - is deliberately left unchanged. Based on v7.3-rc5
(ce1e0223d8ad).

Anlai Lu (6):
iommu/virtio: publish the endpoint before replaying it
iommu/virtio: queue an UNMAP for every removed mapping
iommu/virtio: allocate the UNMAP request with the mapping
iommu/virtio: stop queueing and draining once the device is removed
iommu/virtio: read the endpoint count under the lock in the iotlb
paths
iommu/virtio: batch the mapping replay on domain attach

drivers/iommu/virtio-iommu.c | 642 +++++++++++++++++++++++++++++------
1 file changed, 538 insertions(+), 104 deletions(-)

--
2.55.0