Re: [PATCH v2] mm/mglru: fix and remove redundant unevictable folio handling
From: Ketan Kishore
Date: Wed Aug 26 2026 - 14:12:41 EST
On 8/12/2026 5:52 PM, Kairui Song wrote:
sort_folio() has a shortcut for moving folios that are no longer
evictable but are still sitting on a generation list. However, this
shortcut is buggy. It does not follow the PG_lru usage convention,
and it has a more serious issue.
Unevictable folios are not threaded on lists[LRU_UNEVICTABLE], so that
folio->lru can be reused to hold folio->mlock_count (see the comment in
lruvec_init()). Hence lruvec_add_folio() skips the list_add() for them,
and every other place that turns a folio unevictable initialises
mlock_count explicitly: lru_add() sets it to 0, __mlock_folio() and
__mlock_new_folio() set it to !!folio_test_mlocked(folio).
sort_folio() sets nothing, and the lru_gen_del_folio() right above it
may have already poisoned folio->lru via list_del(), so mlock_count
ends up aliasing LIST_POISON2, which reads as 0x122, i.e. 290. The
result is user visible. On munlock, __munlock_folio() decrements that
bogus count, finds it still non-zero and bails out before clearing
PG_mlocked, so the folio remains unevictable and the Mlocked
accounting stays inflated until the folio is freed.
The shortcut also touches the LRU flags in the wrong order. It calls
lru_gen_del_folio() while PG_lru is still set, so a concurrent
folio_test_clear_lru() (e.g. compaction, folio_isolate_lru()) can
succeed on a folio that has already been taken off the generation list,
which may lead to unexpected behavior.
So fix it by isolating them as common folios and letting the generic
shrink path cull them. This matches the classical LRU behavior, and
there should be no visible effect on the generic eviction or isolation
behavior.
There is no performance concern either, such a folio goes through this
once, and then it is off the generation lists for good.
Fixes: ac35a4902374 ("mm: multi-gen LRU: minimal implementation")
Signed-off-by: Kairui Song <kasong@xxxxxxxxxxx>
We reported an issue with evict_folios() at
https://lore.kernel.org/all/20260807-evict_folios_race-v1-1-b167c6b4cfde@xxxxxxxxxxxxxxxx/
with the following trace:
list_del corruption. prev->next should be fffffffeead4fbc8,
but was ffffeafeead44188. (prev=fffffffee4fc4c08)
kernel BUG at lib/list_debug.c:64!
Call trace:
__list_del_entry_valid_or_report+0x100/0x14c
evict_folios+0x145c/0x16dc
try_to_shrink_lruvec+0x228/0x35c
shrink_one+0x94/0x158
shrink_many+0x1c8/0x1f4
lru_gen_shrink_node+0x94/0x110
shrink_node+0x468/0x8b4
balance_pgdat+0x4f0/0x9a0
kswapd+0x268/0x470
Our v1 fix wrapped the folio_putback_lru() call in evict_folios() with
lruvec->lru_lock. That introduced a self-deadlock: folio_putback_lru()
-> folio_add_lru() can synchronously reach folio_batch_move_lru() ->
folio_lruvec_relock_irqsave().
We had a v2 in progress that instead tracks whether any
unevictable folio was returned via folio_putback_lru() in the lockless
section, and calls lru_add_drain_all() before move_folios_to_lru() only
when needed, to flush all per-CPU LRU-add batches and close the race
window without ever holding lru_lock across folio_putback_lru().
Your patch fixes the same underlying race through a simpler route:
removing the unevictable shortcut in sort_folio() and no longer
special-casing unevictable folios in evict_folios() lets them fall
through to move_folios_to_lru(), which already drops lru_lock before
calling folio_putback_lru() for them. That sidesteps the lock-ordering
hazard entirely and avoids the need for an explicit drain.
Reviewed-by: Ketan Kishore <ketan.kishore@xxxxxxxxxxxxxxxx>
---
Changes in v2:
- Proactively bypass MGLRU pid protection and lazy promotion to avoid
hot unevcitable folios staying on list for a long time.
- Link to v1: https://patch.msgid.link/20260811-mglru-mlock-fix-v1-1-8b2321d0e1d3@xxxxxxxxxxx
---
mm/vmscan.c | 19 +++++--------------
1 file changed, 5 insertions(+), 14 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 3194da7dcc79..ca2b926520ea 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -4648,7 +4648,6 @@ void lru_gen_reparent_memcg(struct mem_cgroup *memcg, struct mem_cgroup *parent,
static bool sort_folio(struct lruvec *lruvec, struct folio *folio, struct scan_control *sc,
int tier_idx)
{
- bool success;
int gen = folio_lru_gen(folio);
int type = folio_is_file_lru(folio);
int zone = folio_zonenum(folio);
@@ -4660,15 +4659,9 @@ static bool sort_folio(struct lruvec *lruvec, struct folio *folio, struct scan_c
VM_WARN_ON_ONCE_FOLIO(gen >= MAX_NR_GENS, folio);
- /* unevictable */
- if (!folio_evictable(folio)) {
- success = lru_gen_del_folio(lruvec, folio, true);
- VM_WARN_ON_ONCE_FOLIO(!success, folio);
- folio_set_unevictable(folio);
- lruvec_add_folio(lruvec, folio);
- __count_vm_events(UNEVICTABLE_PGCULLED, delta);
- return true;
- }
+ /* unevictable: let it through and the generic path will cull it */
+ if (!folio_evictable(folio))
+ return false;
/* promoted */
if (gen != lru_gen_from_seq(lrugen->min_seq[type])) {
@@ -4921,11 +4914,9 @@ static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
list_for_each_entry_safe_reverse(folio, next, &list, lru) {
DEFINE_MIN_SEQ(lruvec);
- if (!folio_evictable(folio)) {
- list_del(&folio->lru);
- folio_putback_lru(folio);
+ /* move_folios_to_lru() culls unevictable folios via folio_putback_lru() */
+ if (!folio_evictable(folio))
continue;
- }
/* retry folios that may have missed folio_rotate_reclaimable() */
if (!skip_retry && !folio_test_active(folio) && !folio_mapped(folio) &&
---
base-commit: 1029098ee3275ea5b78e329ce132262affa2f8cc
change-id: 20260811-mglru-mlock-fix-20d8f8d4847a
Best regards,
--
Kairui Song <kasong@xxxxxxxxxxx>