Re: [PATCH v4 0/4] make unused huge shrinker memcg aware
From: Qi Zheng
Date: Thu Aug 27 2026 - 22:49:04 EST
Hi Andrew,
On 8/28/26 6:50 AM, Andrew Morton wrote:
On Mon, 17 Aug 2026 17:03:24 +0800 Qi Zheng <qi.zheng@xxxxxxxxx> wrote:
Changes in v4:
Changes in v3:
Changes in v2:
Thanks for the diligent versioning info. fyi, it is conventional to
maintain this below the --- separator. It's not really the most
important part of the [0/N]!
Got it.
The shmem unused huge shrinker maintains a per-superblock list of inodes
whose tail huge folio extends beyond i_size. Because this list is not
memcg aware, reclaim triggered by memcg A can scan inodes across the
entire superblock and split huge folios charged to unrelated memcg B,
causing unexpected impact on it.
In the worst case, memcg A has no reclaimable shmem at all, making the
reclaim entirely useless and incurring unnecessary latency. We observed
this in production, where page lock contention during split caused
multi-hundred-millisecond stalls:
Ugh. That's the most important part!
tid 11340 comm scanner locked a page for 182264 us! kstack:
unlock_page+1
split_huge_page_to_list+3135
shmem_unused_huge_shrink+767
super_cache_scan+329
do_shrink_slab+291
shrink_slab+533
shrink_node+400
do_try_to_free_pages+206
try_to_free_mem_cgroup_pages+262
try_charge_memcg+591
mem_cgroup_charge+136
__handle_mm_fault+2431
handle_mm_fault+194
do_user_addr_fault+462
__do_page_fault+176
do_page_fault+48
page_fault+62
Usama's recent patch [1] prevents the shmem unused shrinker from being
invoked during memcg-level reclaim altogether, but this is overly
conservative: we can do better by reclaiming only the shmem charged to
the reclaiming memcg.
This series converts the shrinker list to a memcg-aware list_lru, so
that non-root memcg reclaim walks only candidates charged to the
reclaiming memcg. Global reclaim, root memcg reclaim and shmem quota
reclaim retain their existing global semantics.
To avoid pinning a dying memcg through a long-lived CSS reference, each
inode stores an obj_cgroup reference instead of a mem_cgroup reference.
The list_lru add/delete paths resolve the current memcg from the objcg
under RCU, staying consistent with list_lru's own memcg migration on
offline.
Sashiko said a few things and they look disturbing-if-true:
https://sashiko.dev/#/patchset/cover.1786955972.git.zhengqi.arch@xxxxxxxxxxxxx
(Apologies if this has already been considered - we don't have ways of
tracking all this (yet, I hope) apart from personal memory and personal
memorys are quite fried at present)
As both Usama and I have pointed out [1][2], [PATCH v4 1/4] is the part
that got dropped during the merge. The complete patch [3] was actually
reviewed a while ago.
[1].https://lore.kernel.org/all/20260810101954.822260-1-usama.arif@xxxxxxxxx/
[2]. https://lore.kernel.org/all/9c7efd5f-f8d3-4926-acb4-34c326ffb1c3@xxxxxxxxx/
[3]. https://lore.kernel.org/all/20260715103516.2410175-1-usama.arif@xxxxxxxxx/
Thanks,
Qi