Re: [PATCH v5 01/11] mm, swap: add virtual swap device infrastructure
From: Nhat Pham
Date: Mon Sep 28 2026 - 09:15:23 EST
On Sat, Sep 26, 2026 at 8:02 PM Chris Li <chrisl@xxxxxxxxxx> wrote:
>
> On Fri, Sep 25, 2026 at 8:04 AM Nhat Pham <nphamcs@xxxxxxxxx> wrote:
> >
> > On Thu, Sep 24, 2026 at 5:44 PM Chris Li <chrisl@xxxxxxxxxx> wrote:
> > >
> > > On Thu, Sep 24, 2026 at 6:20 AM Nhat Pham <nphamcs@xxxxxxxxx> wrote:
> > > >
> > > > On Wed, Sep 23, 2026 at 2:18 AM Chris Li <chrisl@xxxxxxxxxx> wrote:
> > > > >
> > > > >
> > > > > There are too many swap_is_vswap() in this series. It fragments the
> > > > > code path, making things harder to reason about. Esepcially around
> > > > > locks. I count 22 in this patch alone.
> > > > >
> > > > > I consider this the biggest drawback of this series. This
> > > > > fragmentation of the code path.
> > > >
> > > > This will mirror my response at [1], but I'm responding here for the
> > > > record and for your convenience.
> > >
> > > Thank you. Really appriciate that.
> > >
> > > > It is needed because vswap *is* a special device, with its own
> > > > requirements. If you don't special case-it, you'd get undesirable
> > > > behaviors.
> > >
> > > Ack, that is why I am against making it a special device. That is the
> > > whole idea behind xswap. It is just normal swap device implemented via
> > > swap_ops just like ext4 is a file system implemnet on top of VFS.
> > >
> > > > Let's take Baoquan's patch series as an example. It treats the xswap
> > > > device as just another swap device, without any special consideration
> > > > for it. Which creates several problem:
> > >
> > > That is a careful design choice. Keep in mind that design choice
> > > usually has pros and cons.
> > >
> > > I will follow your feedback here.
> > >
> > > > 1. At allocation time, cgroups that disable zswap might get an xswap
> > > > slot, because you don't do any check that the device you allocating
> > > > the swap slots for is xswap. At swap_writeout() time, it is *stuck* -
> > > > we already do the unmapping step, so we cannot reclaim the page.
> > > >
> > > > Ironically a physical swapfile backend could have bailed us out here,
> > > > but it's not in that patch series yet :)
> > >
> > > I am not sure I follow. Are you referring to the uncompressible pages
> > > that xswap would have to reject?
> > >
> > > One way to address that is for xswap to just bite the bullet and store
> > > the uncompressible pages within xswap. That way it will stay out of
> > > the reclaim LRU. Zswap also stores some uncompressible pages. The last
> > > time I checked, it was unfortunately not able to serve my case but
> > > perfect for yours.
> >
> > Not even that. Say you have two kind of workloads in the same host:
> > workload enables disk swap but disable zswap (zswap.max = 0), and
> > workload that wants zswap. Let's even ignore writeback for now.
> >
> > You add 2 "swap devices" to the system: xswap and NVME swapfile.
> >
> > At swap allocation time, xswap patch series does not check carefully
> > which swap device is being used, so you might get an xswap slot for
> > the cgroup that disables zswap. You're stuck at swap_writeout() time.
> > You can't undo it anymore, and you don't even support the physical
> > swapfile backend to fallback.
>
> The xswap is to be used with the swap.tiers. So at the swap.tiers it
> selects xswap or SSD swap. Will that solve your problem?
I mean, we're having that problem right now, no? :)
>
> >
> >
> > >
> > > > Vswap both has swapfile backend, AND allows you to bypass it if the
> > > > cgroup disables zswap. Both involves a bit of swap_is_vswap() check,
> > > > but I think it's worth it :)
> > >
> > > If xswap needs the lower tier it can allocate for one. I don't think
> > > xswap should need to have swap_is_xswap() there. Also xswap does not
> > > have all the baggage of zswap.writeback enable interface.
> >
> > But that's my point. As of now, it does not. Unless you merge the
> > other RFC series, which brings it to roughly the same size as what I'm
> > doing here.
>
> It is cleaner because it avoids the dual personality issue. It also
> paves the way to enable writingback from one swap tier to another.
> That is heading to the right directions. The vswap writeback only
> serves zswap, not others.
Writing back from one swap tier to another doesn't exist right now. I
have not even seen a coherent design to support writeback among swap
tiers. You're just delivering pie-in-the-sky.
The RFC you mentioned is not writeback between tiers. If you think it
does, then maybe you should read it again, It's from a very specific
swap device (xswap, with zswap backend), to a swapfile. Here is what
it does: add an array to store indirection (physical swap entry) to
clusters, and update that array when we need to do backend transfer.
Does that ring a bell? That's because it's the exact same mechanism as
vswap (without proper attribution - not even Suggested-by). If
anything, it's actually a bit less coherent, because you can set xswap
to be lower priority than physical swapfile.
Vswap does not have that, because vswap is not an ordinary swap
device. It's implementing the concept of virtualization (the "V" in
VFS, if you want to make that analogy), made optional to reduce
overhead when indirection is not needed. That gives you all the
machinery you need for backend transfer: a centralized virtualized
layer, where you can update the backend, without having to go and
chase down all the external references - no PTE walk needed.
>
> The biggest hurdle is the user API. One of Johannes's biggest
> objections is that it lacks writeback support and perhaps more. Now
> the writeback RFC is available. Let's re-evaluate what is possible.
> The swap.tier series has been out for a year now.
See above. But transferring between tier is still not a solved problem.
>
> If it is the same but does not have the 33 swap_is_vswap() fragmented
> dual personality issue, I see that is a reason to choose that path
> instead. What do you say?
The problem is the interface and the charging question. We have put
out our objections to it, and so far we have not been satisfied with
the defence.
>
> >
> > > >
> > > > Maybe instead of just grepping for swap_is_vswap() and throw your hand
> > > > at the complexity, read the code and try to understand why it's
> > > > necessary first, then propose simplifications if you have good ideas
> > > > in minds.
> > >
> > > Yes, I already read your code I decide that most of the
> > > swap_is_vswap() should not be needed if taking a proper swap_ops
> > > approach. I wanted to offer you the chance to work closely with me to
> > > achieve that, but you did not follow that suggestion.
> > >
> > > Quote from my old email:
> > > https://lore.kernel.org/all/CACePvbX+3tO91BmRwaLqf3Xia32GCemt3W8tayErG81BYLgrzA@xxxxxxxxxxxxxx/
> > >
> > > "I am happy to spend some time working with you to discuss the generic
> > > adopted version of vswap, if you are open to it. Or if you don't want
> > > to waste time on it. I can have someone else or myself come up with
> > > the generic adopted version of vswap for you to review, which I prefer
> > > less."
> > >
> > > Anyway, Baoquan's xswap writeback patch series is out. I suggest
> > > following and improving that series instead.
> > >
> > > > > This is questionable. A vswap device does not go through swap on.
> > > > >
> > > > > This changes the behavior where swap devices always go through swap on/off.
> > > > >
> > > > > I wish the swap device went through swap on to be able to start swapping.
> > > > >
> > >
> > > >
> > > > What usecase would necessitate the need for swapon/swapoff interface?
> > > > Other than just trying to shoehorn it to an existing interface?
> > >
> > > I like universal interfaces. Your vswap will break the /etc/fstab. The
> > > current suggestion for xswap using sysfs will break that as well.
> > > I would like to keep the /etc/fstab interface working, including
> > > support for UUIDs for swap devices, etc.
> >
> > "Universal interfaces" is not a use case. None of the thing you list
> > out here is a use case. It's an interface, but you have not
> > constructed a single use case to justify all of that.
>
> We don't have to go throgh the universal interface to resolve the 33
> swap_is_vswap(). It seems there is an alternative now. We can take the
> xswap route to eventually solve all your problems.
> >
> > > I am very concerned that vswap will start kicking in without me
> > > explicitly enabling it. That changes the previous user-visible
> > > behavior. It breaks my mental model. Please take this feedback
> > > seriously.
> >
> > I mean, there's a knob for it. How can it start kicking in without
> > your enablement?
>
> Please correct me if I am wrong; here is my understanding from you:
>
> vswap is off by default because of that knob. You turn it on and it
> starts zswap.
> Two questions here:
> 1) Is your knob compatible with the existing /etc/fstab where people
> use it to enable swap devices, including with UUIDs, etc.?
If they don't turn it on, their code works as is. Their existing
/etc/fstab exhibits no behavioral change.
What you're asking is us *enrolling* in that same interface. And my
question is what does that interface give us? I already pointed out
multiple ways where that interface's generality shoots the user in the
foot.
> 2) When you enable the knob, do some unit test and disable the knob
> again. Does the vswap go through the same path as swap off so no usage
> of vswap exists on the system?
It's a boottime parameter.
Why would you care about swapoff? What use case does it serve?
>
> That is the existing use case I have in mind.
You're describing an interface, again. It's still not a use case.
I can promise that it won't break /etc/fstab if vswap is off. - if it
does then it's a bug and I can fix it.
>
> Code maintainability is also a consideration when choosing which
> technical solutions to move forward with, when there are alternatives
> to provide solutions to the same problem.
I already pointed out how xswap's design *creates* problem, including
userspace problem. You just handwaved it away with "let the users
shoot themselves in the foot". I disagree with that philosophy.