Re: [RFC PATCH 0/5] iommupt: Introduce IO page table shrinker
From: Pranjal Shrivastava
Date: Mon Oct 05 2026 - 00:59:05 EST
On Fri, Oct 02, 2026 at 12:08:44PM -0300, Jason Gunthorpe wrote:
> On Thu, Oct 01, 2026 at 11:02:14PM +0000, Pranjal Shrivastava wrote:
> > Introduce a lockless, deferred reclamation framework for IOMMU page tables
> > built on the generic_pt library. As VMMs and userspace drivers map and
> > unmap large, sparse IOVA regions through VFIO and iommufd, page table
> > directories are often left allocated but completely empty. generic_pt
> > frees a table when a single unmap covers it entirely, but tables that
> > empty through a series of partial unmaps stay allocated until the domain
> > is destroyed. Under memory pressure, this *stranded* memory cannot be
> > reclaimed and has resulted in OOMs.
> >
> > This series refcounts leaf directories natively in struct ioptdesc and
> > registers a domain-aware MM shrinker that prunes empty directories under
> > system memory pressure.
>
> We had talked about doing it this way
>
> But I had proposed a different, and possibly simpler, solution that
> addresses *just* the iommufd use case.
>
> After unmapping something have iommufd compute the gap in IOVA that
> contains what was unmapped and then issue a 'clean(gap)' operation to
> generic_pt.
>
> This is the same operation as unmap, except we know now that the gap
> has no PTEs so all it does is clean up the table pointers.
>
> This requires no special refcounting or anything difficult beyond
> some locking in iommufd to hold the gap stable while we clean it.
>
> Would it work for you? It seems substantially simpler, but I never
> tried to implement it.
>
I was tempted to use the interval trees too, but I started thinking
about:
a) Locking: For the gap to stay stable while we clean it, we'd need to
add some kind of serialization either through iova_rwsem or a dedicated
gap_lock to prevent a concurrent map to allocate IOVA from that gap.
Thus, every unmap pays for clean under the lock. I haven't perf-ed it
yet, but I'm not sure whether users/Guests using virtio-iommu, where any
guest DMA unmap becomes an unmap on the host, would regress.
b) Other users of IOMMU API (unmanaged domain) like the type 1 (which
was the one hurting our systems) and other in-tree drivers that use an
unmanaged domain and call iommu_map()/iommu_unmap() with their own IOVA
allocator. One of the goals was to avoid enabling the user-space or
in-kernel IOVA allocation (ab)users to cause OOMs via IOPT allocation.
> An alternative version is closer to what you have here, somehow
> connect iommufd to the shrinker and have it lock and walk the gaps
> cleaning them on shrink requests?
>
I like this one better than cleaning on every unmap, since it keeps
the cost off the map/unmap path and only does work under memory
pressure. Although, it doesn't help type1 or the in-tree iommu_map()
users.
Would you consider those users worth covering, or would you rather
they track their own holes and call clean() themselves? Happy to dig
into this at the LPC session too.
Thanks,
Praan