Re: [PATCH v3 00/14] mm, swap: extendable swap devices (xswap)
From: Baoquan He
Date: Mon Sep 21 2026 - 06:22:58 EST
On 09/20/26 at 11:52pm, Chris Li wrote:
> On Thu, Sep 17, 2026 at 3:17 AM Johannes Weiner <hannes@xxxxxxxxxxx> wrote:
> >
> > On Thu, Sep 17, 2026 at 03:31:23PM +0800, Baoquan He wrote:
> > > On 09/16/26 at 12:45pm, Johannes Weiner wrote:
> > > > On Wed, Sep 16, 2026 at 06:19:07PM +0800, Baoquan He wrote:
> > > > > xswap is a swap device with no backing storage. Swapped-out pages live
> > > > > in zswap. Its cluster_info[] array lives in a VM_SPARSE vmalloc area,
> > > > > and the area is grown and shrunk on demand as swap usage changes.
> > > > >
> > > > > The problem being solved is the static size of compressed swap. Both
> > > > > zram and zswap need the size fixed in advance, and neither gives memory
> > > > > back when the workload shrinks. The solution should be a device whose
> > > > > size can scale up/down as per usage. xswap does that by mapping the
> > > > > metadata lazily instead of reserving it for the whole range.
> > > > >
> > > > > Design
> > > > > ------
> > > > > - si->cluster_info[] stays a plain array. Access is still
> > > > > &si->cluster_info[offset / SWAPFILE_CLUSTER]: no per-access branch, no
> > > > > RCU discipline, no tear-down state machine, no NULL return.
> > > > > - Only an initial chunk is mapped at creation. The rest of the address
> > > > > space is reserved, not allocated, so an idle device costs nothing.
> > > > > - Growth is driven by allocation. When no free cluster is left and the
> > > > > address space has room, the next chunk is mapped and added to the free
> > > > > list. No userspace involvement.
> > > > > - Shrink is driven by frees. The free tail is scanned, and whole chunks
> > > > > are unmapped once the mapped range is at most half in use and several
> > > > > chunks can go. One chunk is left mapped as slack, so the next
> > > > > allocation does not map it straight back. A ceiling lowered below the
> > > > > mapped range skips the half-in-use rule and is enforced at once.
> > > >
> > > > If the swap maintainers prefer the VM_SPARSE route, I'm happy to defer
> > > > to them on that.
> > > >
> > > > However, from the cgroup and zswap camp, two stipulations that I
> > > > reasoned out in the other thread[1]:
> > >
> > > >
> > > > 1. You must not charge compression space as swap space to the cgroup.
> > >
>
> Sorry let me push back on that. That is already existing user-visible
> behavior. Changing that will break our deployment using zswap. I don't
> think we should change that.
>
> See more in my reply in the other email thread.
>
> https://lore.kernel.org/linux-mm/CACePvbVaPDnva8X-Xmz84r7j5HjTuih-w5phpw7cerK2uPnK6w@xxxxxxxxxxxxxx/
Thank both for valuable input. I am thinking if we can add an counter
like memory.swap.disk.* or memory.swap.backing.*, then we won't break
the existing behaviour, and also cover the use case Johannes mentioned
where different cgroup have different swapout target setting on xswap.
>
> > > Hmm, I don't have a stance on this. However, isn't this an issue
> > > zswap/zram have been doing? It feels like an independent issue which
> > > should be done separately?
> >
> > If you have 3 containers using compression space, and two of them have
> > writeback enabled to a shared swapfile, the memory.swap.* controls
> > need to work to manage fair access to that swapfile. They do not work
> > if compression space itself is conflated in.
> >
> > Right now zswap entries actually consume physical swapfile space, even
> > before writeback. Charging the space is correct. But the whole point
> > is to decouple compression space from physical swap space.
> >
> > This is not something that can be done later. It would be a dramatic
> > user-visible change to how the resource is categorized and managed.
> >
> > > > 2. You must make the compression space large enough to be outside the
> > > > range where users can hit space limits before hitting memory limits.
> > >
> > > We may need a way to define 'large enough' at first.
> >
> > I've tried to lay this out in the other thread, and highlighted the
> > usability issues that result from hitting compression space limits
> > prematurely. It's kind of your call whether you want to seriously
> > engage with this or not.
> >
> > But ultimately it's your claim that a static size can be made to work,
> > so it's on you to make a convincing case.
> >
> > > > That also means not allowing setups where this is possible.
> > >
> > > And the limit is only an optional knob. If the admin does not set it,
> > > the device grows to the full address space, so there is no space limit
> > > to hit at all. It already behaves the way you want by default. The knob
> > > is only for admins who want a ceiling, they can use it or not. I hope
> > > this would not be a problem for your use case.
> >
> > No, I've laid this out already as well.
> >
> > This isn't about "my" usecase. It's about designing a coherent
> > interface that works well with a large number of usecases, and other
> > pieces of kernel infrastructure commonly used in conjunction.
> >
> > The other proposal in the room needs no such interface. The burden of
> > proof for adding one is on you.
> >
> > > > [1] https://lore.kernel.org/linux-mm/aqLi6cIjD2wJwk0B@xxxxxxxxxxx/