Re: [RFC PATCH] mm: vmscan: avoid anon scanning for GFP_NOIO with low swapcache in MGLRU
From: Barry Song
Date: Sat Oct 10 2026 - 11:22:35 EST
On Sat, Oct 10, 2026 at 10:11 PM Xiang Liu <liuxiang.277@xxxxxxxxxxxxx> wrote:
>
> Bo Zhang's GFP_NOIO patch [1] addressed anonymous scanning with low
> swapcache in the traditional LRU and left MGLRU as a TODO. This patch
> addresses that remaining case.
>
> Under GFP_NOIO, anonymous folios requiring swap IO cannot be reclaimed.
> When very few anonymous folios are already in swapcache, scanning anon
> can spend substantial time isolating and processing folios with little
> reclaim progress.
>
> Apply the existing can_reclaim_anon_pages() check to MGLRU.
> lruvec_is_sizable() excludes unavailable anon from its size estimate.
> In try_to_shrink_lruvec(), file-only scan_swappiness is passed to
> get_nr_to_scan(), should_run_aging() and evict_folios().
>
> The tradeoff is synchronous aging. With four anon generations and two
> file generations, advancing max_seq can require moving the oldest anon
> folios to the next generation. This patch accepts that cost to avoid
> low-yield anonymous scanning. It keeps the original swappiness in
> try_to_inc_max_seq(), avoiding additional use of the inc_min_seq()
> single-type shortcut that advances min_seq without processing the
> oldest folios and can cause cold/hot inversions.
>
> Tested against the mm-unstable base-commit below using dm-verity
> hash-block reads through dm-bufio, with unchanged NOIO allocation flags.
>
> Allocation time spans alloc_buffer() entry to return with valid buffer
> and data. Reclaim time is the time spent in try_to_free_pages() within
> that allocation. Eviction and aging are the time spent in evict_folios()
> and try_to_inc_max_seq() within that reclaim.
>
> 1. NOIO allocations under memory pressure
>
> The workload triggered NOIO reclaim during dm-verity reads, with
> generations allowed to evolve naturally. The table reports median
> times for successful allocations that entered direct reclaim.
>
> Time (ms, median) Baseline Patched
> =========================== ======== =======
> Successful alloc_buffer() 181.196 0.465
> Direct reclaim 181.185 0.131
> Eviction within reclaim 179.147 0.091
>
> 2. Synchronous-aging cost
>
> Prepared four anon generations and two file generations, with zero
> swapcache and pages in the oldest file generation. The tests varied
> the oldest anon population to compare baseline eviction with patched
> synchronous aging.
>
> Each row reports a successful alloc_buffer() call and its reclaim
> costs. Oldest anon is the measured population of the oldest generation.
>
> Kernel Oldest anon Allocation Reclaim Eviction Aging
> (MiB) (ms) (ms) (ms) (ms)
> ======== =========== ========== ======== ======== =======
> Baseline 1755.266 131.704 131.699 127.985 1.672
> Patched 1840.422 38.316 38.311 0.307 37.785
> Baseline 4835.418 681.796 681.766 675.477 4.382
> Patched 4813.180 91.161 91.140 0.450 90.407
> Baseline 8941.582 1474.245 1474.202 1469.084 0.488
> Patched 9012.660 143.748 143.709 1.026 142.341
> Baseline 13104.676 1098.372 1098.354 1087.976 2.307
> Patched 13046.621 139.450 139.390 1.083 137.954
> Baseline 17187.043 2694.909 2694.902 2686.470 0.784
> Patched 17124.902 209.240 209.233 0.638 208.270
>
> Across these tested sizes, synchronous aging was substantially cheaper
> than baseline eviction, and successful allocation time decreased.
>
> [1]
> https://patchew.org/linux/20260908062649.1045883-1-zhangbo56@xxxxxxxxxx/
>
> Signed-off-by: Xiang Liu <liuxiang.277@xxxxxxxxxxxxx>
> ---
> mm/vmscan.c | 26 ++++++++++++++++----------
> 1 file changed, 16 insertions(+), 10 deletions(-)
>
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index 29b31446fc11..550c16cad9f9 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -379,13 +379,6 @@ static inline bool reclaimable_anon_is_low(struct mem_cgroup *memcg,
> if (!sc || (sc->gfp_mask & __GFP_IO))
> return false;
>
> - /*
> - * FIXME: MGLRU doesn't fully respect can_reclaim_anon_pages() for the
> - * scanning type, so only apply this to the traditional LRU for now.
> - */
> - if (lru_gen_enabled())
> - return false;
> -
> if (memcg) {
> struct lruvec *lruvec = mem_cgroup_lruvec(memcg, pgdat);
>
> @@ -4357,6 +4350,13 @@ static bool lruvec_is_sizable(struct lruvec *lruvec, struct scan_control *sc)
> unsigned long total;
> int swappiness = get_swappiness(lruvec, sc);
> struct mem_cgroup *memcg = lruvec_memcg(lruvec);
> + int nid = lruvec_pgdat(lruvec)->node_id;
> +
> + if (swappiness && !can_reclaim_anon_pages(memcg, nid, sc)) {
> + if (swappiness == SWAPPINESS_ANON_ONLY)
> + return false;
> + swappiness = MIN_SWAPPINESS;
> + }
Hi Xiang,
Thanks for your patch.
I'm not a fan of this fix, as setting
scan_swappiness = MIN_SWAPPINESS will force a cold/hot
inversion for anonymous folios.
static bool inc_min_seq(struct lruvec *lruvec, int type, int
swappiness)
{
...
/* For file type, skip the check if swappiness is anon only */
if (type && (swappiness == SWAPPINESS_ANON_ONLY))
goto done;
/* For anon type, skip the check if swappiness is zero (file
only) */
if (!type && !swappiness)
goto done;
...
}
We should try to avoid this.
Best Regards
Barry