Re: [PATCH 0/8] Optimize anonymous swapbacked large folio unmapping
From: David Hildenbrand (Arm)
Date: Wed Aug 19 2026 - 03:45:13 EST
On 8/19/26 08:50, Dev Jain wrote:
>
>
> On 23/07/26 12:38 pm, Dev Jain wrote:
>> Speed up unmapping of anonymous swapbacked large folios by clearing
>> the ptes, and setting swap ptes, in one go.
>>
>> The following benchmark (stolen from Barry) is used to measure the
>> time taken to swapout 256M worth of memory backed by 64K large folios:
>>
>> #define _GNU_SOURCE
>> #include <stdio.h>
>> #include <stdlib.h>
>> #include <sys/mman.h>
>> #include <string.h>
>> #include <time.h>
>> #include <unistd.h>
>> #include <errno.h>
>>
>> #define SIZE_MB 256
>> #define SIZE_BYTES (SIZE_MB * 1024 * 1024)
>>
>> int main() {
>> void *addr = mmap(NULL, SIZE_BYTES, PROT_READ | PROT_WRITE,
>> MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
>> if (addr == MAP_FAILED) {
>> perror("mmap failed");
>> return 1;
>> }
>>
>> memset(addr, 0, SIZE_BYTES);
>>
>> struct timespec start, end;
>> clock_gettime(CLOCK_MONOTONIC, &start);
>>
>> if (madvise(addr, SIZE_BYTES, MADV_PAGEOUT) != 0) {
>> perror("madvise(MADV_PAGEOUT) failed");
>> munmap(addr, SIZE_BYTES);
>> return 1;
>> }
>>
>> clock_gettime(CLOCK_MONOTONIC, &end);
>>
>> long duration_ns = (end.tv_sec - start.tv_sec) * 1e9 +
>> (end.tv_nsec - start.tv_nsec);
>> printf("madvise(MADV_PAGEOUT) took %ld ns (%.3f ms)\n",
>> duration_ns, duration_ns / 1e6);
>>
>> munmap(addr, SIZE_BYTES);
>> return 0;
>> }
>>
>> Performance as measured on a Linux VM on Apple M3 (arm64):
>>
>> Vanilla - Mean: 37401913 ns, std dev: 12%
>> Patched - Mean: 17420282 ns, std dev: 11%
>>
>> resulting in more than 2x speedup.
>>
>> No regression observed on 4K folios.
>>
>> Performance as measured on bare metal x86:
>>
>> Vanilla - mean: 54986286 ns, std dev: 1.5%
>> Patched - mean: 51930795 ns, std dev: 3%
>>
>> I tried magnifying the difference on x86 by using 1M large folios, but
>> can't spot an obvious improvement (looks like my system is too fast to
>> benefit from batched atomic operations!), hinting that the benefit lies
>> mainly in the reduction of ptep_get() calls and the reduction of TLB
>> flushes during contpte-unfolding, on arm64.
>>
>> No regression is observed on 4K folios on x86 too.
>>
>> ---
>
>
> Hi Andrew,
>
> I am not sure what is the current process in mm regarding sending patches during
> merge window - would you prefer seeing v2 after merge window or shall I respin?
>
>
Let's wait for more review first. Most of the patches have not been reviewed and
there is no time to rush at this point.
--
Cheers,
David