RE: [PATCH] mm/vma: don't remove VMA from rmap if pgoff unchanged
From: Deng, Pan
Date: Tue Sep 29 2026 - 09:51:26 EST
> > > The anonymous rmap is keyed on anon_vma_chains not VMAs, so in those
> > > instances anon_rmap_tree_update_vma_inplace() iterates over
> > > vma->anon_vma_chain, invoking anon_rmap_tree_update_inplace() on
> each one.
> > >
> > > For the anon rmap case, with CONFIG_DEBUG_VM_RB set, avc-
> >cached_vma_last
> > > is also updated in anon_rmap_tree_update_inplace().
> > >
> > > When performing a VMA shrink or a split where the VMA is the lower one,
> the
> > > page offset cannot change, so set the flags unconditionally in these cases.
> > >
> > > When merging VMAs the page offset is unchanged only in some cases, so
> > > update init_multi_vma_prep() to set the flags only if the page offsets
> > > remain the same.
> > >
> > > These changes ultimately result in less rmap lock contention.
> >
> > I think this asks for numbers?
>
> Well I don't have any :)
>
> It logically reduces the contention, and that can only be a good thing.
>
> Pan had some numbers, I've asked him to re-run against this one.
>
Yes, UnixBench/execl case, measured on my 2-socket, 192 core / 384
Thread x86-64 system, on v7.3-rc4 with and without this patch, built
from an identical kernel .config.
Configuration, re-applied after each boot:
- cpufreq governor "performance"
- uncore frequency pinned to max
Run rule: 10 runs per kernel, 30s cool-down in between, cmd:
$ ./Run execl -c 384
Result: Execl Throughput, index score:
avg %stdev min max
v7.3-rc4 3511.5 0.44% 3494.5 3543.4
+ patch 4069.0 0.48% 4047.4 4109.5
The speedup is +15.9%.
In addition, I also profiled lock contention data ~10s in the
middle of one iteration, cmd:
$ sudo perf lock contention -ab -l -S vma_prepare -E 8
The dominant lock is i_mmap_rwsem
Result: avg wait on that lock in ms, 5 runs per kernel:
avg %stdev min max
v7.3-rc4 9.470 2.91% 9.070 9.820
+ patch 8.420 3.34% 8.100 8.770
avg wait is ~11.1% reduction.
Note: under the i_mmap_rwsem write lock, this case only exercises the
split path, for both the file and the anon rmap, while merge and shrink
are not reached when the lock is held.
Best Regards
Pan