Re: [PATCH v4 10/13] mm/collapse: open-code collapse_single_pmd() in its two callers
From: Kiryl Shutsemau
Date: Fri Oct 02 2026 - 10:28:05 EST
On Thu, Oct 01, 2026 at 11:37:31AM +0200, David Hildenbrand (Arm) wrote:
> On 9/28/26 12:06, Kiryl Shutsemau wrote:
> > From: "Kiryl Shutsemau (Meta)" <kas@xxxxxxxxxx>
> >
> > A scan and a collapse want different things from mmap_lock. The scan
> > reads one PTE table under the lock the caller holds, refuses most of the
> > time, and the caller moves on to the next table without letting go. The
> > collapse allocates, may sleep in writeback and takes the lock for write
> > itself, so the lock it is handed is of no use to it.
> >
> > collapse_single_pmd() kept that boundary inside itself. It dropped the
> > lock on some paths and not others, and reported which by way of a bool
> > its callers had to carry along and then act on.
>
> Let me think this through. collapse_scan_file: it doesn't actually need the MM at
> all. The only reason is to do tracing. Rather stupid, it just should not consume
> the MM at all anymore. Consequently it doesn't even need the mmap lock. But the
> caller needs the mmap lock to figure out the file + range from the vma (the
> per-vma lock would also be sufficient for that).
Right on both.
As I mentioned before, virtual address scanning might not be the best
way to collapse page cache.
We might want to start from inodes on superblocks that have large folios
enabled.
But it is out of scope for the patchset.
> Also, I guess we can convert some of the scanning to use per-vma locks in the
> future, whereby we would actually want to scan with the per-vma lock held.
That is the reason the caller owns the lock here.
On top of this series I have khugepaged scanning under the per-VMA lock:
the scan asserts whatever lock the caller holds, the caller does
vma_end_read() before the run, and the run takes the locks it needs
itself.
Neither the scan nor the run changed for it.
With collapse_single_pmd() dropping the lock, it has to know which lock
the caller took. process_madvise() on another process's mm cannot move
to the VMA lock: untagged_addr_remote() asserts mmap_lock, so a remote
MADV_COLLAPSE keeps it while a local one and khugepaged move on.
The single call would need a flag saying which lock to drop. That is
the lock_dropped bool again, pointing the other way.
> I do wonder about one thing: should we really care so much about keeping the
> mmap lock locked? Meaning, why not provide a single collapse_single_pmd() that
>
> * Is always called without the mmap lock (as is)
> * Always returns with the mmap lock unlocked (change)
Yes, we should.
Walking the tables under one hold and giving the lock up only to
collapse is how khugepaged has worked since ba76149f47d8 ("thp:
khugepaged").
The bool is just how that signal got plumbed when MADV_COLLAPSE arrived
in 50ad2f24b3b4, and it has already cost one bug, 5a62019807da
("mm/khugepaged: fix issue with tracking lock").
This series removes the bool and keeps the behaviour.
> Sure, we drop+re-acquire the mmap lock a couple of times and lookup the vma, but
> isn't that actually being nice to the other parts of the system? In the future
> it would simply get called with the per-vma lock and would return with it unlocked.
>
> In the good old days, looking up VMAs was expensive, but nowadays ... not sure
> if it still matters?
The rwsem and the VMA lookup are not the only cost.
In your version every table is a visit to the mm: unlock,
khugepaged_mm_lock, back through khugepaged_do_scan(),
khugepaged_mm_lock again, a trylock and a VMA walk, for a scan that on
memory already huge is one pmd read.
I measured your prototype against patch 9 as posted, both on a production
config, 8G of anonymous memory already collapsed, scan_sleep_millisecs 0,
pages_to_scan 65536, PMD order only, five runs each, medians:
as posted yours
khugepaged, nothing to collapse:
full passes over 8G in 20s 251,180 55,477
khugepaged CPU 100% 100%
cost per refused table 19 ns 88 ns
MADV_COLLAPSE, already huge:
512M per call, p50 3.8 us 17.4 us
2M per call, p50 703 ns 703 ns
tables walked in full then refused
(max_ptes_none 0), passes in 20s 803 836
On being nice to the rest of the system: the scan is already bounded.
The read hold ends after pages_to_scan worth of progress. We already
have properly sized scan batching in place.
> khugepaged? Not sure if this matters. madvise? I suspect many real users operate
> on a single PMD only (e.g., tcmalloc, jemalloc). For the other ones, not sure if
> dropping the lock every PMD is really a problem?
khugepaged matters. We run it across the fleet, and its scan cost is
what decides how aggressive we can afford to be with it. I am not
giving up scan rate to keep one entry point into the engine.
For a single PMD there is no difference either way. For a range, today a
MADV_COLLAPSE over memory that is already huge walks it without letting
go of the lock at all; in your version it relocks and revalidates for
every table, and tells madvise_walk_vmas() to look the VMA up again.
> IOW, how bad would the following simplification be (prototype that needs more
> work and thought):
>
...
>
> Based on that, I'd rather want to see collapse_single_pmd() to just inline the file
> and anon paths, and see how we can further optimize the locking internally (e.g., perform
> the pagecache scanning without the mmap lock).
I see the appeal of the single entry point for the collapse engine. I do.
But it costs us on both the locking picture and the scan rate. That
does not work for me.
--
Kiryl Shutsemau / Kirill A. Shutemov