Re: Path forward for Virtualized Swap?

From: Nhat Pham

Date: Mon Sep 28 2026 - 08:39:51 EST


On Fri, Sep 25, 2026 at 9:26 PM Kairui Song <ryncsn@xxxxxxxxx> wrote:
>
> On Fri, Sep 25, 2026 at 09:54:25AM +0800, Nhat Pham wrote:
> > On Fri, Sep 25, 2026 at 6:15 AM Kairui Song <ryncsn@xxxxxxxxx> wrote:
> > >
> > > On Tue, Sep 22, 2026 at 09:44:37PM +0100, Chris Li wrote:
> > > > It seems you are talking about a different topic: the vswap charging issue.
> > > > There is a golden rule that we should follow: don't break existing
> > > > users. At least with the same persistence, this rule should apply
> > > > universally.
> > > > In the swap tiers discussion, the UAPI was such a big deal that we
> > > > couldn't implement new UAPI. On the other hand here we argue for
> > > > liberally changing user-space visible behavior.
> > > >
> > > > BTW, I already shared that changing swap counter charging will break
> > > > our and others' existing deployments.
> > >
> > > Hi all
> >
> > Hi Kairui,
> >
> > Thank you for your thoughtful response! Lots of food for thought for me :)
>
> Hi Nhat, thanks for the reply!
>
> > >
> > > Just for reference. Maybe a seperate counter, tiering, is a better idea
> > > than changing the swap counter?
> >
> > I think memory.swap.* is never meant to be used as the "offloaded
> > footprint". It is incidentally correct, because of architectural
> > limitations: zswap/swap cache/zero page usage leads to real
> > consumption of real resources (physical swapfile).
>
> It's not that "incidentally" I think? TGhe defination from the function
> level seems pretty clear, folio_alloc_swap -> charge. A logical
> limitation.
>

I don't think so. It's just that until now, architecturally, physical
swap is always required, regardless of the state a swapped out page is
in (whether it's in swap cache, in zswap, or already wrriten out to
disk). So folio_alloc_swap() gets you a physical swap slot, which you
need to charge. That changes when a "swap slot" can now be virtual.

memcg folks can explain better than I can, but cgroupv2 is meant to
isolate concrete physical resources, because physical resources are
often limited/static, and can therefore be contended. As Johannes
commented, that's sort of why memsw was not ported to cgroupv2 - it's
not an actual "physical" resource, but rather just a logical quantity
(that's useful to reason about). You can also find this quote in the
documentation that explicitly spells out that cgroupv2 is meant for
physical resources:

"For trusted jobs, on the other hand, a combined counter is not an
intuitive userspace interface, and it flies in the face of the idea
that cgroup controllers should account and limit specific physical
resources. Swap space is a resource like all others in the system, and
that’s why unified hierarchy allows distributing it separately."

So I don't think memory.swap is meant to represent "offloaded
footprint". I mean even in the current state, that conceptualization
is already shaky - a page in swap cache is resident in memory. If you
mincore() that page, it's reported as resident. It's not offloaded.
But it's counted towards memory.swap because memory.swap tracks
physical swap usage, and swap cache page, despite being resident,
still occupies physical swap (until vswap is in the picture).

A really good analogy would be virtual vs physical memory. Correct me
if I'm wrong, but I believe we do not charge/limit virtual memory
usage at the cgroup level (at least in v2)?

I think it is fair to expose useful counters that users can monitor and act
upon. We have a bunch of such counters (anon/file usage, thp usage,
etc). But virtual swap usage is not a resource that can create
contention or isolation issue independently of the backend it take. I
don't think it should be handled/charged similar to (or together with)
physical swap!