Re: [PATCH] mm/list_lru: don't copy stale shrinker id from non-memcg-aware shrinkers
From: Andrew Morton
Date: Tue Sep 01 2026 - 13:42:36 EST
On Tue, 1 Sep 2026 19:51:04 +0800 Qinyun Tan <qinyuntan@xxxxxxxxxxxxxxxxx> wrote:
> With cgroup.memory=nokmem, shrinker_memcg_alloc() fails with -ENOSYS
> for shrinkers without SHRINKER_NONSLAB, and shrinker_alloc() falls
> back to a non-memcg-aware shrinker. On this fallback path,
> shrinker->id is never assigned and keeps 0 from kzalloc(), which is a
> valid id belonging to whichever memcg-aware shrinker registers first.
>
> __list_lru_init() copies shrinker->id unconditionally, so every
> list_lru backed by such a fallback shrinker (thp-deferred_split,
> zswap-shrinker, workingset shadow nodes, superblock lrus, ...) ends
> up with lru->shrinker_id == 0 instead of -1.
>
> Under nokmem the list_lru collapses to the shared per-node lists, but
> __list_lru_add() still calls set_shrinker_bit() against the memcg of
> the added object. Most list_lru users are unaffected because their
> objects resolve to a NULL memcg without kmem accounting, but the THP
> deferred split queue holds user folios, which are charged regardless
> of nokmem. Since no memcg-aware shrinker can register under nokmem,
> shrinker_nr_max stays 0 and every memcg's shrinker_info has
> map_nr_max == 0, so the first folio added by khugepaged triggers on
> every boot:
>
> WARNING: mm/shrinker.c:212 at set_shrinker_bit+0x99/0xa0
>
> On systems where a SHRINKER_NONSLAB shrinker (btrfs, xfs) did register
> and expand the maps, there is no warning; instead bit 0 is set
> spuriously for an unrelated shrinker.
>
> shrinker->id is only meaningful while SHRINKER_MEMCG_AWARE is set,
> and all readers inside mm/shrinker.c already check the flag before
> using the id. Make __list_lru_init() do the same and fall back to -1,
> so set_shrinker_bit() is never reached with a bogus id. The stale
> shrinker->id itself is left as is; cleaning that up is a separate
> topic.
Thanks. I'll queue this for test and review.
> Fixes: 03375203e1da8 ("mm: do not allocate shrinker info with cgroup.memory=nokmem")
Worth a cc:stable, I assume.
> ---
>
> Verified on a machine booting with cgroup.memory=nokmem and
> CONFIG_TRANSPARENT_HUGEPAGE=y: the warning fires once per boot from
> khugepaged, disappears when nokmem is removed from the command line,
> and no longer triggers with this fix applied and nokmem set.
That's useful info. I'll move it into the changelog.
Sashiko thinks there's a problem with cgroup_disable=memory as well:
https://sashiko.dev/#/patchset/20260901115104.2944996-1-qinyuntan@xxxxxxxxxxxxxxxxx