Re: [PATCH v2 0/5] read proc/pid/smaps_rollup under per-vma lock
From: Suren Baghdasaryan
Date: Tue Sep 08 2026 - 12:25:30 EST
On Tue, Sep 8, 2026 at 9:05 AM Xueyuan Chen <xueyuan.chen21@xxxxxxxxx> wrote:
>
>
> On Sun, Sep 06, 2026 at 11:39:13PM -0700, Suren Baghdasaryan wrote:
> >proc/pid/smaps_rollup can be read using the combination of RCU and
> >VMA read locks, similar to proc/pid/{maps|smaps|numa_maps}. RCU is
> >required to safely traverse the VMA tree and VMA lock stabilizes the
> >VMA being processed and the pagetable walk.
> >Note that we have to keep the logic to drop mmap_lock on contention
> >because even when using per-VMA locks we might have to fall back to
> >holding the mmap_lock.
> >
>
> Hi Suren,
>
> I tested this series on my arm64 24-core machine, using the same
> command:
>
> run-proc-vs-map.sh --nsamples 20 --rawdata -- \
> --busyduration 2 --procfile smaps_rollup
>
> baseline:
> Median Minimum Maximum
> 0.132 0.121 2.600
> 0.132 0.123 2.768
> 0.132 0.123 2.804
> 0.132 0.125 2.835
> 0.132 0.125 2.837
> 0.132 0.125 2.849
> 0.132 0.125 2.858
> 0.132 0.125 2.888
> 0.132 0.130 2.894
> 0.132 0.132 2.896
> 0.132 0.132 2.915
> 0.132 0.134 2.921
> 0.132 0.134 2.927
> 0.132 0.137 2.945
> 0.132 0.141 2.984
> 0.132 0.246 3.004
> 0.132 0.270 3.381
> 0.132 1.316 3.533
> 0.132 1.321 3.551
> 0.132 1.343 3.924
>
> patched:
> Median Minimum Maximum
> 0.006 0.006 1.950
> 0.006 0.006 1.955
> 0.006 0.006 1.955
> 0.006 0.006 1.956
> 0.006 0.006 1.956
> 0.006 0.006 1.957
> 0.006 0.006 1.958
> 0.006 0.006 1.959
> 0.006 0.006 1.959
> 0.006 0.006 1.961
> 0.006 0.006 1.962
> 0.006 0.006 1.963
> 0.006 0.006 1.964
> 0.006 0.006 1.966
> 0.006 0.006 1.969
> 0.006 0.006 1.976
> 0.006 0.006 1.982
> 0.006 0.006 1.995
> 0.006 0.006 1.996
> 0.006 0.006 2.043
>
> So the median is 0.132 -> 0.006 ms, about 22x, and the worst case
> drops from 3.9 to 2.0 ms.
>
> Tested-by: Xueyuan Chen <xueyuan.chen21@xxxxxxxxx>
Thanks for the confirmation!
>
> Thanks,
> Xueyuan
>
> >The first 3 patches are cleanups making later change simpler. The main
> >change is in patch 4. Patch 5 extends existing proc-maps-race tearing
> >test to verify smaps_rollup content.
> >
> >Changes since v1 [1]
> >- removed helper functions which are not needed after per-VMA locks
> >became unconditional in [2];
> >- added a cover letter, per Lorenzo Stoakes and David Hildenbrand;
> >- removed is_mmap_lock_contended() as it's no more needed after [2],
> >per David Hildenbrand;
> >- fixed typos, per Lorenzo Stoakes and David Hildenbrand;
> >- refactored duplicate error checks, per Lorenzo Stoakes;
> >- eliminated special case of start=0 in smap_gather_stats(),
> >per Lorenzo Stoakes;
> >- updated comments, per Lorenzo Stoakes;
> >- moved original comment about 4 possible cases when mmap lock is
> >dropped into patch 4 changelog;
> >- added smaps_rollup testing into proc-maps-race test.
> >
> >[1] https://lore.kernel.org/all/20260606015729.1837935-1-surenb@xxxxxxxxxx/
> >[2] https://lore.kernel.org/all/20260813193433.3318288-1-surenb@xxxxxxxxxx/
> >
> >Suren Baghdasaryan (5):
> > proc/task_mmu: remove unnecessary helpers
> > proc/task_mmu: remove unnecessary inlines in function definitions
> > proc/task_mmu: remove special-casing of smap_gather_stats() start
> > parameter
> > proc/task_mmu: read proc/pid/smaps_rollup under per-vma lock
> > selftests/proc: add /proc/pid/smaps_rollup tearing tests
> >
> > fs/proc/task_mmu.c | 280 +++++++-----------
> > tools/testing/selftests/proc/proc-maps-race.c | 187 +++++++++++-
> > 2 files changed, 291 insertions(+), 176 deletions(-)
> >
> >
> >base-commit: e3fc12b08aadde9cec7b3799ac0e0c9a1aa245c4
> >--
> >2.55.0.979.g7e5102b832-goog
> >
> >