Re: [PATCH v2 3/5] proc/task_mmu: remove special-casing of smap_gather_stats() start parameter
From: Suren Baghdasaryan
Date: Thu Sep 10 2026 - 19:34:57 EST
On Thu, Sep 10, 2026 at 9:27 AM Lorenzo Stoakes (ARM) <ljs@xxxxxxxxxx> wrote:
>
> On Wed, Sep 09, 2026 at 09:16:23PM +0200, David Hildenbrand (Arm) wrote:
> > On 9/9/26 20:28, Suren Baghdasaryan wrote:
> > > On Wed, Sep 9, 2026 at 10:16 AM David Hildenbrand (Arm)
> > > <david@xxxxxxxxxx> wrote:
> > >>
> > >> On 9/7/26 08:39, Suren Baghdasaryan wrote:
> > >>> smap_gather_stats() interprets its start parameter to mean vma->vm_start
> > >>> when it's set to 0. Eliminate this special interpretation and pass
> > >>> vma->vm_start explicitly when needed.
> > >>>
> > >>> Since smap_gather_stats() operates within a single VMA, we can replace
> > >>> walk_page_vma()/walk_page_range() calls with walk_page_range_vma()
> > >>> which is simpler and also can be called while holding per-VMA lock.
> > >>>
> > >>> No functional change intended.
> > >>>
> > >>> Suggested by: Lorenzo Stoakes <ljs@xxxxxxxxxx>
> > >>> Signed-off-by: Suren Baghdasaryan <surenb@xxxxxxxxxx>
> > >>> ---
> > >>> fs/proc/task_mmu.c | 40 ++++++++++++++++++++++------------------
> > >>> 1 file changed, 22 insertions(+), 18 deletions(-)
> > >>>
> > >>> diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
> > >>> index 9908ba32f180..3351decd1172 100644
> > >>> --- a/fs/proc/task_mmu.c
> > >>> +++ b/fs/proc/task_mmu.c
> > >>> @@ -1246,20 +1246,27 @@ get_smaps_shmem_walk_ops(struct proc_maps_private *priv)
> > >>> return &smaps_shmem_walk_vma_lock_ops;
> > >>> }
> > >>>
> > >>> -/*
> > >>> - * Gather mem stats from @vma with the indicated beginning
> > >>> - * address @start, and keep them in @mss.
> > >>> +/**
> > >>> + * smap_gather_stats() - Gather mem stats from @vma.
> > >>> + * @priv: proc maps private state.
> > >>> + * @vma: The VMA to gather stats for.
> > >>> + * @mss: The accumulated stats.
> > >>> + * @start: The address from which to start.
> > >>> *
> > >>> - * Use vm_start of @vma as the beginning address if @start is 0.
> > >>> + * This gathers stats for the whole of the VMA unless the lock was dropped
> > >>> + * and VMA grew or got merged and we found it again, in which case we only
> > >>> + * gather stats for the remainder of the VMA range.
> > >>> */
> > >>> static void smap_gather_stats(struct proc_maps_private *priv,
> > >>> struct vm_area_struct *vma,
> > >>> - struct mem_size_stats *mss, unsigned long start)
> > >>> + struct mem_size_stats *mss,
> > >>> + unsigned long start)
> > >>> {
> > >>> const struct mm_walk_ops *ops = get_smaps_walk_ops(priv);
> > >>> + const bool is_partial = start > vma->vm_start;
> > >>>
> > >>> /* Invalid start */
> > >>> - if (start >= vma->vm_end)
> > >>> + if (start < vma->vm_start || start >= vma->vm_end)
> > >>> return;
> > >>>
> > >>> if (vma == get_gate_vma(priv->lock_ctx.mm))
> > >>> @@ -1279,20 +1286,17 @@ static void smap_gather_stats(struct proc_maps_private *priv,
> > >>> * Unless we know that the shmem object (or the part mapped by
> > >>> * our VMA) has no swapped out pages at all.
> > >>> */
> > >>> - unsigned long shmem_swapped = shmem_swap_usage(vma);
> > >>> + const unsigned long shmem_swapped = shmem_swap_usage(vma);
> > >>> + const bool shared_or_ro = vma_test(vma, VMA_SHARED_BIT) ||
> > >>> + !vma_test(vma, VMA_WRITE_BIT);
> > >>>
> > >>> - if (!start && (!shmem_swapped || (vma->vm_flags & VM_SHARED) ||
> > >>> - !(vma->vm_flags & VM_WRITE))) {
> > >>> + if (!is_partial && (!shmem_swapped || shared_or_ro))
> > >>> mss->swap += shmem_swapped;
> > >>> - } else {
> > >>> + else
> > >>> ops = get_smaps_shmem_walk_ops(priv);
> > >>> - }
> > >>
> > >> Horrible, horrible code, really. But not your fault :)
> > >>
> > >> I think we can just make the shared_or_ro less odd by just checking for cow
> > >> mappings (as described in the comment).
> > >>
> > >> const bool is_cow = vma_is_cow_mapping(vma);
> > >>
> > >> ...
> > >>
> > >> if (is_partial || (shmem_swapped && is_cow))
> > >> ops = get_smaps_shmem_walk_ops(priv);
> > >> else
> > >> mss->swap += shmem_swapped;
> > >>
> > >> That's almost in a form that I could understand what's happening.
> > >
> > > Hmm. So, are you saying that !is_cow always implies shared_or_ro? Or
> > > maybe you are stating that vma_is_cow_mapping() was the actual intent
> > > here?
> >
> > So the comment says:
> >
> > "For private writable mappings, we might have COW pages that .."
> >
> > Which translates to:
> >
> > private writable == vma_is_cow_mapping()
> >
> > >
> > > shared_or_ro = VMA_SHARED_BIT || !VMA_WRITE_BIT
> > >
> > > is_cow = !VMA_SHARED_BIT && VMA_MAYWRITE_BIT
> > > !is_cow = VMA_SHARED_BIT || !VMA_MAYWRITE_BIT
> > >
> > > so, !is_cow would impy shared_or_ro only if !VMA_MAYWRITE_BIT always
> > > implies !VMA_WRITE_BIT. But I think it's possible to have a VMA that
> > > has VMA_WRITE_BIT but not VMA_MAYWRITE_BIT, right?
> >
> > VMA_WRITE should imply VMA_MAYWRITE
>
> Yup it's illegal to set VMA_WRITE_BIT without VMA_MAYWRITE_BIT.
>
> >
> > (in sanitize_fault_flags() we even disallow write faults entirely if VM_MAYWRITE
> > is missing)
> >
> > For example, a
> > > driver can create such a VMA to allow writing to the VMA but to lock
> > > its content once mprotect(PROT_READ) gets called.
>
> I'm not even sure how a driver would achieve that but drivers in general are not
> permitted to alter VMA flags after map time.
>
> >
> > I don't think that would be valid for a driver to do. But it wouldn't matter
> > here because
> >
> > shmem_mapping(vma->vm_file->f_mapping)
>
> Yup :)
>
> >
> >
> > I think we could simplify the comment as well to:
> >
> > diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
> > index e671b4fd8dedd..4b7e7089cafa7 100644
> > --- a/fs/proc/task_mmu.c
> > +++ b/fs/proc/task_mmu.c
> > @@ -1303,12 +1303,10 @@ static void smap_gather_stats(struct proc_maps_private
> > *priv,
> >
> > if (vma->vm_file && shmem_mapping(vma->vm_file->f_mapping)) {
> > /*
> > - * For shared or readonly shmem mappings we know that all
> > - * swapped out pages belong to the shmem object, and we can
> > - * obtain the swap value much more efficiently. For private
> > - * writable mappings, we might have COW pages that are
> > - * not affected by the parent swapped out pages of the shmem
> > - * object, so we have to distinguish them during the page walk.
> > + * In CoW mappings, we might have anon folios that are
> > + * independent of the shmem object. So fallback to the less
> > + * efficient mechanism in such mappings.
> > + *
>
> Maybe tweak to 'CoW mappings might map anon folios that do not belong to shmem,
> so perform a less efficient page table walk in this situation' or something like
> that?
I ended up with this:
/*
* CoW mappings might map anon folios that do not belong to
* shmem. Perform a less efficient page table walk in this
* situation, unless we know that the shmem object (or the
* part mapped by our VMA) has no swapped out pages at all.
*/
Hope this explains the condition clearly.
>
>
> > * Unless we know that the shmem object (or the part mapped by
> > * our VMA) has no swapped out pages at all.
> > */
>
>
>
> >
> > --
> > Cheers,
> >
> > David
>
> --
> Cheers, Lorenzo