Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()

From: Yosry Ahmed

Date: Thu Jul 23 2026 - 12:49:21 EST


On Thu, Jul 23, 2026 at 6:55 AM Johannes Weiner <hannes@xxxxxxxxxxx> wrote:
>
> On Wed, Jul 22, 2026 at 09:52:18PM -0700, Yosry Ahmed wrote:
> > On Wed, Jul 22, 2026 at 7:27 PM Johannes Weiner <hannes@xxxxxxxxxxx> wrote:
> > >
> > > On Fri, Jul 17, 2026 at 04:51:51PM +0800, Hao Jia wrote:
> > > > @@ -1369,7 +1402,7 @@ static void shrink_worker(struct work_struct *w)
> > > > goto resched;
> > > > }
> > > >
> > > > - ret = shrink_memcg(memcg);
> > > > + ret = shrink_memcg(memcg, NR_ZSWAP_WB_BATCH);
> > > > /* drop the extra reference */
> > > > mem_cgroup_put(memcg);
> > > >
> > > > @@ -1493,7 +1526,7 @@ bool zswap_store(struct folio *folio)
> > > > objcg = get_obj_cgroup_from_folio(folio);
> > > > if (objcg && !obj_cgroup_may_zswap(objcg)) {
> > > > memcg = get_mem_cgroup_from_objcg(objcg);
> > > > - if (shrink_memcg(memcg)) {
> > > > + if (shrink_memcg(memcg, 1)) {
> > >
> > > Why 64 for the global limit but only 1 for the cgroup limit? That
> > > seems arbitrary in multiple ways.
> >
> > I suggested that we keep the writeback here without batching and do
> > that change separately, mainly out of abundance of caution as
> > writeback is done synchronously here so the extra latency could be
> > problematic. I think we probably want to measure the performance
> > impact of that separately.
> >
> > That being said, this path is potentially too expensive anyway due to
> > the flush, but I would rather we do some basic measurements before
> > batching here.
> >
> > What do you think?
>
> It's not an unknown, right? We know this works for direct reclaimers,
> cgroup limit reclaim e.g., and what the latency implications are.
>
> Because of how reclaim works, we also know it'll call zswap_store() in
> batches of SWAP_CLUSTER_MAX. If we don't batch here, they're likely to
> each call shrink_memcg() once we're at the limit - while still risking
> rejections due to compressibility differences.
>
> My worry is that if we start with an inconsistency, we'll be stuck
> with it for a long time.
>
> I'd rather start with the clean, consistent version. Dial it back only
> if we have data to justfiy the complication that we can put into a
> comment and the changelog that outlines why exactly it's different.

I am fine with doing that and basically always using NR_ZSWAP_WB_BATCH
as the batch size in shrink_memcg(), but I would be more comfortable
if we did some sanity testing.

Hao, would you be able to do some smoke testing with NR_ZSWAP_WB_BATCH
used for all paths, and memory.zswap.max set in a way that induces
writeback? You can probably set memory.zswap.max to 1% of total memory
instead of the global pool limit and rerun the same test.