Re: [RFC PATCH RESEND v5 0/2] mm/vmscan: fix NR_ISOLATED accounting and throttle MGLRU eviction
From: Barry Song
Date: Wed Sep 23 2026 - 17:55:15 EST
On Fri, Sep 18, 2026 at 4:07 PM Hui Zhu <hui.zhu@xxxxxxxxx> wrote:
>
> From: Hui Zhu <zhuhui@xxxxxxxxxx>
>
> Previously the patch was based on mm-unstable, which caused issues during
> Sashiko apply.
> So I rebased it onto mm-stable and resend.
> Thanks Andrew for the heads-up.
>
> The legacy reclaim path has two mechanisms around isolated folios that
> MGLRU lacks:
>
> 1. NR_ISOLATED_ANON/FILE counters are updated when folios are isolated
> from the inactive lists. Compaction's too_many_isolated() relies on
> them to decide when to back off. The MGLRU reclaim path never
> updates them, so compaction cannot see MGLRU's in-flight isolation.
>
> 2. shrink_inactive_list() throttles direct reclaim via
> too_many_isolated() when isolated folios pile up. MGLRU's
> evict_folios() isolates folios without any such check, so many
> concurrent reclaimers can over-isolate the same (oldest) generation,
> leading to unnecessary swapping, thrashing and premature memcg OOM.
>
> Patch 1 fixes the counter accounting in the MGLRU isolation path.
> Patch 2 adds a per-lruvec throttle, mirroring the legacy behavior but
> adapted to MGLRU's per-lruvec contention and its dynamic type
> selection/fallback.
>
> Testing
> =======
> Two test scripts are provided to reproduce the problem and validate
> the fix. Both are available at:
> https://gist.github.com/teawater/3ef51251f2e91a5a600e3d26bb477e34
> Test environment: 10 CPU / 8GB QEMU guest, MGLRU enabled, a 16MB
> memory cgroup, anonymous working set, swap backed by dm-delay (50ms
> read/write delay) to slow swap-out and lengthen the isolation window.
>
> mglru_iso_repro.sh (64 threads, 48MB working set, 60s):
> Drives concurrent direct reclaim inside the memcg and measures scan
> efficiency, throttle events, in-flight isolation and throughput.
> Neither kernel OOMs at this concurrency; the value of the patch shows
> in reclaim quality:
> unpatched patched
> OOM kills 0 0
> mm_vmscan_throttled 0 35645 (all
> VMSCAN_THROTTLE_ISOLATED)
> nr_isolated peak 0 (invisible) 230
> total touches 246,499,132,369 301,456,426,692
> scan efficiency 0.0261 0.0194
>
> The patched kernel completes ~22% more work in the same 60s: the
> throttle keeps concurrent reclaimers from trampling the same
> generation, so less CPU is burned in reclaim. Note nr_isolated is
> always 0 on the unpatched kernel - the over-isolation is invisible
> there, which is exactly what patch 1 fixes.
>
> mglru_iso_repro_v2.sh (192 threads, 48MB working set, 60s):
> Raises concurrency to the point where over-isolation becomes fatal:
> unpatched patched
> OOM kills 1 (task killed) 0
> memcg oom events 51 0
> mm_vmscan_throttled 0 512292 (all
> VMSCAN_THROTTLE_ISOLATED)
> total touches 0 (killed) 682,300,003,972
>
> With 192 threads the unpatched kernel cannot keep reclaim ahead of
> allocation and the task is OOM-killed; the patched kernel survives the
> full run and keeps reclaim making progress.
>
Hi Hui,
I’m perfectly fine with using a script to create a scenario with a small
memcg (e.g. 16 MB or 64 MB) and many threads. This is how we do R&D and
build simple reproducers for complex problems.
But I’m still struggling to understand how this scenario maps to or
affects real-world workloads, other than `iso_stress.c` in your scripts.
Do you have any idea of a real workload that suffers from the lack of
throttling, which we could present in the cover letter?
Best Regards
Barry