Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()

From: Nhat Pham

Date: Fri Jul 17 2026 - 12:46:59 EST


On Fri, Jul 17, 2026 at 1:52 AM Hao Jia <jiahao.kernel@xxxxxxxxx> wrote:
>
> From: Hao Jia <jiahao1@xxxxxxxxxxx>
>
> Currently, shrink_memcg() writes back at most one entry per-node during
> its traversal. This makes shrink_worker() inefficient, as it must
> repeatedly re-enter shrink_memcg() to make any substantial progress.
> Under high memory pressure, this can cause the writeback speed to be
> too slow to keep up with refaults, leading to zswap store failures and
> forcing pages to skip zswap and go directly to disk, which results in
> an LRU inversion.
>
> To address this, extend shrink_memcg() and rewrite its LRU iteration logic
> to support batch writeback. Introduce the nr_to_scan parameter to bound how
> many pages are scanned per call. This enables batch writeback in the
> shrink_worker() path, while maintaining a low scan budget in the
> zswap_store() path.
>
> Test Setup:
> - Total memory: 32 GB.
> - zswap settings: max_pool_percent=1, accept_threshold_percent=50,
> shrinker_enabled=N.
>
> Test Case 1:
> Allocate 512MB of anonymous pages and fill them with random data (to avoid
> compression), then use cgroup memory.reclaim to force a large amount of
> anonymous pages into zswap. At an interval of 2ms, allocate a 4K anonymous
> page where the first 4 bytes are random numbers and the rest are zeros, and
> then trigger a reclamation of this 4K anonymous page through cgroup
> memory.reclaim. When the pool threshold is reached, shrink_memcg() will
> be triggered.
> The test data after running for 120s is as follows:
> Baseline Patched
> shrink_worker wakeups 5,363 85
> shrink_memcg calls 11,373,201 180,928
> written_back pages 40,212 40,236
> zswap_store calls 161,190 168,741
> store succeeded (ret=1) 102,743 127,644
> store rejected (ret=0) 58,447 41,097
> store reject rate ~36% ~24%
> pool_limit_hit delta 55,826 14,062
> pswpout 98,659 81,333
> pswpin 2 1
>
> Test Case 2:
> To consistently force zswap store failures and trigger shrink_worker(),
> the following stress-ng command was run for 120 seconds within a cgroup
> limited to a memory.max of 1G:
> bash -c 'echo $$ > /sys/fs/cgroup/zswaptest/cgroup.procs ; \
> exec stress-ng --vm 4 --vm-bytes 4G --vm-keep --vm-method rand-set -t \
> 120s -q'
> The test data after running for 120s is as follows:
> Baseline Patched
> shrink_worker wakeups 5,640 987
> shrink_memcg calls 8,481,500 2,504,818
> written_back pages 260 768,576
> zswap_store calls 2,742,756 2,301,414
> store succeeded (ret=1) 934,640 1,308,686
> store rejected (ret=0) 1,808,116 992,728
> store reject rate ~66% ~43%
> pool_limit_hit delta 1,181,310 101,593
> pswpout 1,808,376 1,761,304
> pswpin 4,288,497 3,902,658
>

Acked-by: Nhat Pham <nphamcs@xxxxxxxxx>