Re: Path forward for Virtualized Swap?
From: Kairui Song
Date: Fri Sep 25 2026 - 15:27:25 EST
On Fri, Sep 25, 2026 at 11:41:31AM +0800, Johannes Weiner wrote:
>
> Thanks for your thoughtful email, Kairui.
Hello Johannes,
> > I also want to separate two things that I think got bundled together
> > here: not requiring a physical slot behind a compressed entry, and not
> > charging the raw size to the swap counter. The first one is great, yeah,
> > and it's exactly the part we want, it's what makes compression usable
> > without provisioning disk. The second one is a policy change, maybe it's
> > not needed for the first stage, charging a cgroup for the
> > memories it has offloaded doesn't require any slot to exist behind them.
> > If someone wants to run memory compression with no disk at all,
> > memory.swap.max defaults to max, so that still works fine, right?
>
> I think what we found out over the course of this discussion is that
> people have been using memory.swap.max in two ways. Regardless of what
> we do, we will "break" one side.
>
> (1) The usecase you're describing. Use memory.swap.max, combined with
> memory.max, to set a "total", predictable footprint of in-use
> application address space. You can mmap whatever you want, but the
> number of unique pages you can touch is limited to this sum. And you
> can control residency vs non-residency through the invididual values
> of those settings. If compressed entries are not included, this
> usecase will break.
Right, thanks for the reply! residency vs non-residency is one of
the issues here. It's also about how we consider these two kind of
resources to balance the scheduling of containers.
> (2) The use case we have. Use memory.swap.max to divide a finite space
> in storage. We only have so much space on disk, and we need to manage
> fair access. Note that this isn't about speed. We have a mix of
> containers where some use writeback and others do not. The ones who
Same for us, the usage is mixed.
> write back to the swapfile need to be able to get their fair share -
> not more, not less. Including something that doesn't actually consume
Is that writeback compression rate based? I mean for zswap, you have to
writeback uncompressable part, and then also do cold writeback through
shrinker. The compressable part is hard to predict and control though?
> this separate finite resource is also a behavioral change that would
> break that usecase. And arguably it's a deviation from how the control
> was intended and from the broader cgroup design philosophy.
>
> I think we need to build tools to support both cases. But somebody
> will have to change how they're doing things... :/
>
In the long term we might need something like tier or xxx_ext for
cgroup for better control of different resources? In either case I
think if we could get the building blocks more flexible, and this
actually sounds like a tiering issue? I remember Youngjun once
mentioned a memcg X tier limit which sound a bit similiar.