Re: [PATCH] mm/vma: don't remove VMA from rmap if pgoff unchanged

From: Lorenzo Stoakes (ARM)

Date: Tue Sep 29 2026 - 10:08:36 EST


+cc Barry for anon question below

On Tue, Sep 29, 2026 at 01:45:48PM +0000, Deng, Pan wrote:
> > > > The anonymous rmap is keyed on anon_vma_chains not VMAs, so in those
> > > > instances anon_rmap_tree_update_vma_inplace() iterates over
> > > > vma->anon_vma_chain, invoking anon_rmap_tree_update_inplace() on
> > each one.
> > > >
> > > > For the anon rmap case, with CONFIG_DEBUG_VM_RB set, avc-
> > >cached_vma_last
> > > > is also updated in anon_rmap_tree_update_inplace().
> > > >
> > > > When performing a VMA shrink or a split where the VMA is the lower one,
> > the
> > > > page offset cannot change, so set the flags unconditionally in these cases.
> > > >
> > > > When merging VMAs the page offset is unchanged only in some cases, so
> > > > update init_multi_vma_prep() to set the flags only if the page offsets
> > > > remain the same.
> > > >
> > > > These changes ultimately result in less rmap lock contention.
> > >
> > > I think this asks for numbers?
> >
> > Well I don't have any :)
> >
> > It logically reduces the contention, and that can only be a good thing.
> >
> > Pan had some numbers, I've asked him to re-run against this one.
> >
>
> Yes, UnixBench/execl case, measured on my 2-socket, 192 core / 384
> Thread x86-64 system, on v7.3-rc4 with and without this patch, built
> from an identical kernel .config.
>
> Configuration, re-applied after each boot:
> - cpufreq governor "performance"
> - uncore frequency pinned to max
>
> Run rule: 10 runs per kernel, 30s cool-down in between, cmd:
> $ ./Run execl -c 384
>
> Result: Execl Throughput, index score:
>
> avg %stdev min max
> v7.3-rc4 3511.5 0.44% 3494.5 3543.4
> + patch 4069.0 0.48% 4047.4 4109.5
>
> The speedup is +15.9%.
>
> In addition, I also profiled lock contention data ~10s in the
> middle of one iteration, cmd:
> $ sudo perf lock contention -ab -l -S vma_prepare -E 8
>
> The dominant lock is i_mmap_rwsem
>
> Result: avg wait on that lock in ms, 5 runs per kernel:
>
> avg %stdev min max
> v7.3-rc4 9.470 2.91% 9.070 9.820
> + patch 8.420 3.34% 8.100 8.770
>
> avg wait is ~11.1% reduction.
>
> Note: under the i_mmap_rwsem write lock, this case only exercises the
> split path, for both the file and the anon rmap, while merge and shrink
> are not reached when the lock is held.

Thanks Peng! Much appreciated.

I'll fold this into the commit message on respin, should get that out today
(some trivial renaming, etc. functionality will be identical).

Be interesting to look at anon also, there could be impact for android
specifically given zygote.

Suren - do you have any bandwidth for checking whether this patch impacts anon
rmap lock contention?

Barry - I know you've looked at this in the past, are you able to assess whether
this has impact for your workloads?

>
> Best Regards
> Pan
>

--
Cheers, Lorenzo