Re: [PATCH RESEND 0/9] selftests/mm: improve mremap_test

From: Sarthak Sharma

Date: Tue Sep 29 2026 - 05:14:18 EST


Hi Kalesh and David!

On 9/29/26 1:17 PM, David Hildenbrand (Arm) wrote:
> On 9/29/26 09:32, Kalesh Singh wrote:
>> On Mon, Sep 28, 2026 at 10:54 PM Sarthak Sharma <sarthak.sharma@xxxxxxx> wrote:
>>>
>>>
>>>
>>> On 9/24/26 10:30 AM, Sarthak Sharma wrote:
>>>> This series fixes several correctness issues in mremap_test and
>>>> simplifies and strengthens its data validation.
>>>>
>>>> Patch 1 converts the mremap_test to use kselftest helpers, removes
>>>> manual tracking of failed tests and corrects some spelling errors.
>>>>
>>>> Patches 2 to 5 fix userfaultfd skipping, unexpected mremap success
>>>> handling, multi-VMA data validation and failure reporting when data
>>>> corruption is detected.
>>>>
>>>> Patch 6 removes the randomization and uses a simple pattern based approach.
>>>>
>>>> Patch 7 removes perf tests and timing infrastructure.
>>>>
>>>> Patch 8 removes validation threshold and always validates complete mappings.
>>>>
>>>> Patch 9 strengthens the multi VMA validation by also checking the
>>>> mapping state of holes after remapping.
>>>
>>> Hello everyone! Just wanted to check if someone has had a chance to look
>>> at the series.
>>>
>>> Also, Sashiko has a concern [1], and I had the same while posting the
>>> series. I hope I can get some opinion from the community.
>>>
>>> Currently, for PUD remap tests, we allocate a source mapping of 2GB. We
>>> only fault in the first threshold_mb amount of memory, remap the whole
>>> region, and validate the already faulted threshold_mb amount of memory,
>>> which is, by default, equal to 4MB and can be changed by the user by
>>> supplying a command line option.
>>>
>>> Since we plan to remove all command line options from teh selftests, I
>>> removed threshold_mb altogether, following some discussion on the list
>>> [2]. This would now cause a 2GB source mapping, and its contents copied
>>> from another 2GB buffer. This causes the process to have a 4GB RSS.
>
> Yeah, that's a lot for small CI systems indeed.
>
>>> Sashiko says that this can cause OOM killing in small CI machines.
>>>
>>> Would it be okay to keep a 4GB RSS in this case, or should we find some
>>> other way of validating a part of the whole range instead?
>>
>> Hi Sarthak,
>>
>> IIRC when I initially introduced the test, John was concerned that
>> validating the whole range would significantly increase the duration
>> of the mm selftests; this is why the threshold was introduced. Please
>> check how much it increases if we validate the full range (with
>> David's suggestions) and if it's no longer a concern from other folks.
>> I am fine with removing the threshold.

I tested on an Orion O6. Here's the difference before and after removing
the threshold right now.

Before:
0.02user 0.07system 0:00.10elapsed 98%CPU (0avgtext+0avgdata
87332maxresident)k
0inputs+0outputs (0major+21133minor)pagefaults 0swaps

After:
1.67user 4.91system 0:06.63elapsed 99%CPU (0avgtext+0avgdata
4195564maxresident)k
0inputs+0outputs (0major+2225058minor)pagefaults 0swaps

I think instead of the time taken, the more concerning thing here is RSS
(85MB vs 4GB).

>
> As discussed off-list, I guess it makes more sense to validate a couple of pages
> at the beginning, the middle and the end?

Yup, makes sense. I will implement this then.