Re: [PATCH 0/4] mm/ksm: use per-VMA locking for find_mergeable_vma()
From: Jinjiang Tu
Date: Sun Sep 13 2026 - 23:04:12 EST
在 2026/9/13 12:17, xu.xin16@xxxxxxxxxx 写道:
Nice catch. Thanks for pointing this out. Indeed, the original exclusion betweenPatch 1 removes an unused 'vma' member from struct folio_walk. It hasHi.
never been used since its introduction and is pure cleanup.
Patch 2 adds a 'walk_lock' member to struct folio_walk and extends
folio_walk_start() to assert the required locking mode. Existing
callers are converted to pass PGWALK_RDLOCK, so there is no functional
change. This prepares folio_walk_start() for callers that hold a
per-VMA read lock instead of mmap_read_lock(), which is needed by the
Patch 4. No functional change.
Patch 3 tranforms the boolean 'lock_vma' into the enum 'page_walk_lock'
without any behavior changed, which is prepared for the Patch 4 to use
per-VMA locking. No functional change.
Patch 4 introduces find_mergeable_vma_locked(), which uses the
universal per-VMA locking helper vma_start_read_unlocked() to look up
and read-lock a VM_MERGEABLE VMA without taking mmap_read_lock(). All
KSM call sites that previously used find_mergeable_vma() under
mmap_read_lock() are converted to the new helper, and the locking in
get_mergeable_page() is switched to PGWALK_VMA_RDLOCK_VERIFY so that
folio_walk_start() can verify the per-VMA lock is held.
A microbenchmark was run to measure the time KSM takes to merge a
victim region under mmap_lock contention. Under interference from 4 churner
threads, the merge time of the per-VMA KSM-optimized kernel is
significantly reduced by 50%.
During task exiting, __ksm_exit() uses mmap_write_lock() to synchronize with ksmd.
see the comment of ksm_test_exit().
void __ksm_exit(struct mm_struct *mm)
{
...
if (easy_to_free) {
mm_slot_free(mm_slot_cache, mm_slot);
mm_flags_clear(MMF_VM_MERGE_ANY, mm);
mm_flags_clear(MMF_VM_MERGEABLE, mm);
mmdrop(mm);
} else if (mm_slot) {
mmap_write_lock(mm);
mmap_write_unlock(mm);
}
}
When ksmd currently is scanning the exiting mm, we should guarantee the mm pagetable
still valid (i.e., mm_users > 0). However, ksm_mm_slot only holds mm_count, which only
guarantees the mm_strcut isn't freed. So, __ksm_exit() uses mmap write lock to synchronize
with ksmd.
IIUC, vma_read_lock cannot be exclusive with mmap_write_lock().
__ksm_exit() and ksmd relied on mmap_write_lock() blocking mmap_read_lock(),
and per-VMA read locks do not provide that exclusion.
A possible approach to restore the necessary guarantee is to pin mm_users while
ksmd is walking the page tables:
Before scanning a given mm, try to take a reference with mmget_not_zero(mm).
If it fails, the mm is exiting, so we skip it.
Hold that reference for the entire duration of scanning that mm (not per-VMA),
and drop it with mmput() when done.
On the fallback path where we need to acquire mmap_read_lock(), drop the mm_users
reference before waiting, to avoid delaying an exiting mm.
We should avoid holding mm_users ref too long. Otherwise, when the task being scanned
by ksmd is OOM-skilled, even though the victim task has responsed SIGKILL signal and
exited, the mmaps aren't released due to mm_user > 0.
We should check whether the mm_user has dropped to 1 during scanning, like what
ksm_test_exit() has done.
This directly guarantees that mm_users > 0 while ksmd is accessing the page tables,
so __mmput() cannot reach exit_mmap() and free them. It is more precise than the
old mmap_write_lock() synchronization and should not introduce noticeable delay,
since the reference is only held for the scan duration.