Re: [PATCH 0/4] mm/ksm: use per-VMA locking for find_mergeable_vma()

From: xu.xin16

Date: Wed Sep 23 2026 - 23:41:48 EST


>>> A possible approach to restore the necessary guarantee is to pin
>>> mm_users while
>>> ksmd is walking the page tables:
>>>
>>> Before scanning a given mm, try to take a reference with
>>> mmget_not_zero(mm).
>>> If it fails, the mm is exiting, so we skip it.
>>>
>>> Hold that reference for the entire duration of scanning that mm (not
>>> per-VMA),
>>> and drop it with mmput() when done.
>>>
>>> On the fallback path where we need to acquire mmap_read_lock(), drop
>>> the mm_users
>>> reference before waiting, to avoid delaying an exiting mm.
>>
>> We should avoid holding mm_users ref too long. Otherwise, when the
>> task being scanned
>> by ksmd is OOM-skilled, even though the victim task has responsed
>> SIGKILL signal and
>> exited, the mmaps aren't released due to mm_user > 0.
>
>In this case, we have to relying on OOM reaper to work. But it need to wait OOM_REAPER_DELAY
>(2s) to response.

It shouldn't be too long if holding mm_users only during ksmd accesses mm's pagetable.
Essentially, the delayed time of releasing mmap has no difference to the original approach by
mmap_write_lock.

Or If we want to keep the code simpler, we can consider the other way to keep serialization
and synchronization between __ksm_exit() and ksmd thread, that is: Adding vma_start_write()
for each vma? like:

diff --git a/mm/ksm.c b/mm/ksm.c
index 624f37975e12..1f11646bb00d 100644
--- a/mm/ksm.c
+++ b/mm/ksm.c
@@ -3129,6 +3129,10 @@ void __ksm_exit(struct mm_struct *mm)
mmdrop(mm);
} else if (mm_slot) {
mmap_write_lock(mm);
+ VMA_ITERATOR(vmi, mm, 0);
+ for_each_vma(vmi, vma)
+ vma_start_write(vma);
+
mmap_write_unlock(mm);
}