Re: Path forward for Virtualized Swap?

From: Kairui Song

Date: Fri Sep 25 2026 - 16:26:33 EST


On Fri, Sep 25, 2026 at 09:54:25AM +0800, Nhat Pham wrote:
> On Fri, Sep 25, 2026 at 6:15 AM Kairui Song <ryncsn@xxxxxxxxx> wrote:
> >
> > On Tue, Sep 22, 2026 at 09:44:37PM +0100, Chris Li wrote:
> > > It seems you are talking about a different topic: the vswap charging issue.
> > > There is a golden rule that we should follow: don't break existing
> > > users. At least with the same persistence, this rule should apply
> > > universally.
> > > In the swap tiers discussion, the UAPI was such a big deal that we
> > > couldn't implement new UAPI. On the other hand here we argue for
> > > liberally changing user-space visible behavior.
> > >
> > > BTW, I already shared that changing swap counter charging will break
> > > our and others' existing deployments.
> >
> > Hi all
>
> Hi Kairui,
>
> Thank you for your thoughtful response! Lots of food for thought for me :)

Hi Nhat, thanks for the reply!

> >
> > Just for reference. Maybe a seperate counter, tiering, is a better idea
> > than changing the swap counter?
>
> I think memory.swap.* is never meant to be used as the "offloaded
> footprint". It is incidentally correct, because of architectural
> limitations: zswap/swap cache/zero page usage leads to real
> consumption of real resources (physical swapfile).

It's not that "incidentally" I think? TGhe defination from the function
level seems pretty clear, folio_alloc_swap -> charge. A logical
limitation.

>
> But I understand your concerns regarding the missing observability of
> the "logical" footprint (i.e the "offloaded size"), and how that would
> hamper legitimate use cases.
>
> How about a vswap counter that tracks the vswap usage? That would
> provide the "logical" view. Userspace can then monitor and act based
> on it.
>
> I think that should cover both of the situations that you (and
> Johannes in [1]) pointed out:
>
> (1): With memory.current and this vswap counter, you can kill any
> workloads that violate the level of overcommitting that you set out.
>
> (2): For us, we can use memory.swap.* counter to provide isolation to
> the physical swap space, which is a real, static resource that can be
> hogged.
>
> There will unfortunately be changes in userspace programs/scripts that
> is required. It's unavoidable. We're adding a new behavior. It's the

Right, but usually I think adding new things are generically considered
better than breaking existing things, unless the existing things are
either just metric, or really broken. But memory.swap seems has a pretty
clean definition and implementation, and usage?