Re: Path forward for Virtualized Swap?
From: Chris Li
Date: Sat Sep 26 2026 - 17:28:14 EST
On Fri, Sep 25, 2026 at 5:41 AM Johannes Weiner <hannes@xxxxxxxxxxx> wrote:
>
> On Fri, Sep 25, 2026 at 09:15:17PM +0800, Kairui Song wrote:
> > On Tue, Sep 22, 2026 at 09:44:37PM +0100, Chris Li wrote:
> > > It seems you are talking about a different topic: the vswap charging issue.
> > > There is a golden rule that we should follow: don't break existing
> > > users. At least with the same persistence, this rule should apply
> > > universally.
> > > In the swap tiers discussion, the UAPI was such a big deal that we
> > > couldn't implement new UAPI. On the other hand here we argue for
> > > liberally changing user-space visible behavior.
> > >
> > > BTW, I already shared that changing swap counter charging will break
> > > our and others' existing deployments.
> >
> > Hi all
> >
> > As I read the threads and try to clean up the requirement, just
> > realized that I forgot and ignored something previously. I think I can
> > share a few things here.
> >
> > I also see that Rik mentioning that:
> >
> > > It adds up anonymous, file, accounted slab, and
> > > (after compression) zswap memory use for a cgroup,
> > > and can be limited with all the usual cgroup
> > > limits.
> >
> > That's very true, and that's also the one reason we can't
> > migrate some workload to zswap easily (at least yet) :D
> > See below.
> >
> > The discussion on this can be saw two years before (I know
> > things are different for V2, so see below):
> > https://lore.kernel.org/linux-mm/CAMgjq7AYA91f4g-bknUZOMg6hApTD-X5LqjcTBN2u-Lu8pjs+w@xxxxxxxxxxxxxx/
> >
> > An minor update for that, memsw in V1 serves pretty well (we also
> > modded that part and would try push to upstream if doable), and as
> > memsw is missing in V2, we can still workaround that using
> > memory.current and memory.swap.current. BUt missing the offloaded
> > part in memory.swap seems a problem.
> >
> > First a little bit off topic, I'll be really happy if we can make
> > both compressed memory and swap as separate counters (I even once
> > tried to implement a zpool accounting to account compressed memory
> > in some unified way, but, well, zpool got killed before I post
> > that :P), or at least a way to do that, e.g. something like nokmem.
>
> Thanks for your thoughtful email, Kairui.
>
> > Due to our real usage:
> >
> > With compressed memory staying in a separate counter (which
> > we manged to do that with ZRAM) the memory.current + memory.swap
> > (or, memsw for cgv1) could be the exactly planned or sold size of a
> > container, the scheduler (e.g. from k8s level) is fully aware of
> > the packing rate of a host based on this reading. and can make
> > scheduling decisions based on that. And can control it by
> > adjusting the limit two combined.
> >
> > But with compression as a fixed part in memory.current, first the
> > compression rate is totally uncontrollable, both the user and us
> > will be fully *unaware* of how much memory they can *actually* use,
> > that makes the planning really awkward. memory.max stops being the
> > bound of what we planned or sold, anything compressed lets the raw
> > footprint go past it by however much the compression ratio happens
> > to give, so what we oversold is bounded by the workload's data and
> > not by anything we configure. We can substract the zswap reading
> > though with adaption, however it's hard to change the performance,
> > OOM behavior or reclaim behavior:
> >
> > As you may considering compression is trading CPU time with memory,
> > then two things here: the user could use more memory than we expected
> > by burning the CPU. And, some users has a leaking application, the
> > application could goes super slow or experiencing high CPU usage due
> > to memory being compressed. They really just want to get OOM killed
> > in time when ever the application leaks beyound a threshold (and
> > yes that is a real and actually practical model for many applications).
> > And, we can't simply disable memory compression for them.
> >
> > In many cases we just want a best effort compression to make space
> > for low priority tasks, and do not want ordinary containers to use
> > compression at the cost of lose of performance. While still has
> > a fixed limit as usual. So simply disable memory compression is also
> > not the plan, we do need compression to make place for other
> > applications, we just don't want their real raw usage to exceed
> > memory.max, and we can dynamically adjust memory.swap.max to
> > control the oversold part, compression or physical.
> >
> > And this is not about residency, so memory.min/low don't help here:
> > it's about overselling, and about not leaving a container thrashing in
> > compress/decompress loops instead of being killed.
> >
> > And if the memory compression is really fully transparent (not
> > doable by software), yeah, that's great as there is nothing to do
> > with reclaim. But, for now, we have to go through page fault / folio
> > allocation / map it again. So For example, if we already have
> > memory.max == memory.current or under high pressure, then now
> > doing any read from the compressed part would need to some
> > require further eviction first to make place for the decompressed
> > new data, this is not like any kind of "real" memory, something
> > feels not right here.
> >
> > Another thing is that I think we has been assuming that physical
> > swap is slower than compressed memory, which is not always true either.
> > They all need to be read through page fault, the page fault could
> > be the real blocker here rather than IO or de-compression.
> >
> > I also want to separate two things that I think got bundled together
> > here: not requiring a physical slot behind a compressed entry, and not
> > charging the raw size to the swap counter. The first one is great, yeah,
> > and it's exactly the part we want, it's what makes compression usable
> > without provisioning disk. The second one is a policy change, maybe it's
> > not needed for the first stage, charging a cgroup for the
> > memories it has offloaded doesn't require any slot to exist behind them.
> > If someone wants to run memory compression with no disk at all,
> > memory.swap.max defaults to max, so that still works fine, right?
>
> I think what we found out over the course of this discussion is that
> people have been using memory.swap.max in two ways. Regardless of what
> we do, we will "break" one side.
>
> (1) The usecase you're describing. Use memory.swap.max, combined with
> memory.max, to set a "total", predictable footprint of in-use
> application address space. You can mmap whatever you want, but the
> number of unique pages you can touch is limited to this sum. And you
> can control residency vs non-residency through the invididual values
> of those settings. If compressed entries are not included, this
> usecase will break.
Thanks for the very nice write-up; that is the exact usage case I have
in mind. You put it up there much nicer than I can. Really appreciate
that.
>
> (2) The use case we have. Use memory.swap.max to divide a finite space
> in storage. We only have so much space on disk, and we need to manage
> fair access. Note that this isn't about speed. We have a mix of
> containers where some use writeback and others do not. The ones who
> write back to the swapfile need to be able to get their fair share -
> not more, not less. Including something that doesn't actually consume
Just want to make sure I understand correctly. Do you mean the fair
share of compressed memory used in the zswap case?
> this separate finite resource is also a behavioral change that would
> break that usecase. And arguably it's a deviation from how the control
> was intended and from the broader cgroup design philosophy.
>
> I think we need to build tools to support both cases. But somebody
> will have to change how they're doing things... :/
Agree.
Chris