Re: [PATCH v3 00/14] mm, swap: extendable swap devices (xswap)
From: Chris Li
Date: Mon Sep 21 2026 - 06:56:18 EST
On Thu, Sep 17, 2026 at 3:17 AM Johannes Weiner <hannes@xxxxxxxxxxx> wrote:
>
> On Thu, Sep 17, 2026 at 03:31:23PM +0800, Baoquan He wrote:
> > On 09/16/26 at 12:45pm, Johannes Weiner wrote:
> > > On Wed, Sep 16, 2026 at 06:19:07PM +0800, Baoquan He wrote:
> > > > xswap is a swap device with no backing storage. Swapped-out pages live
> > > > in zswap. Its cluster_info[] array lives in a VM_SPARSE vmalloc area,
> > > > and the area is grown and shrunk on demand as swap usage changes.
> > > >
> > > > The problem being solved is the static size of compressed swap. Both
> > > > zram and zswap need the size fixed in advance, and neither gives memory
> > > > back when the workload shrinks. The solution should be a device whose
> > > > size can scale up/down as per usage. xswap does that by mapping the
> > > > metadata lazily instead of reserving it for the whole range.
> > > >
> > > > Design
> > > > ------
> > > > - si->cluster_info[] stays a plain array. Access is still
> > > > &si->cluster_info[offset / SWAPFILE_CLUSTER]: no per-access branch, no
> > > > RCU discipline, no tear-down state machine, no NULL return.
> > > > - Only an initial chunk is mapped at creation. The rest of the address
> > > > space is reserved, not allocated, so an idle device costs nothing.
> > > > - Growth is driven by allocation. When no free cluster is left and the
> > > > address space has room, the next chunk is mapped and added to the free
> > > > list. No userspace involvement.
> > > > - Shrink is driven by frees. The free tail is scanned, and whole chunks
> > > > are unmapped once the mapped range is at most half in use and several
> > > > chunks can go. One chunk is left mapped as slack, so the next
> > > > allocation does not map it straight back. A ceiling lowered below the
> > > > mapped range skips the half-in-use rule and is enforced at once.
> > >
> > > If the swap maintainers prefer the VM_SPARSE route, I'm happy to defer
> > > to them on that.
> > >
> > > However, from the cgroup and zswap camp, two stipulations that I
> > > reasoned out in the other thread[1]:
> >
> > >
> > > 1. You must not charge compression space as swap space to the cgroup.
> >
Sorry let me push back on that. That is already existing user-visible
behavior. Changing that will break our deployment using zswap. I don't
think we should change that.
See more in my reply in the other email thread.
https://lore.kernel.org/linux-mm/CACePvbVaPDnva8X-Xmz84r7j5HjTuih-w5phpw7cerK2uPnK6w@xxxxxxxxxxxxxx/
Chris
> > Hmm, I don't have a stance on this. However, isn't this an issue
> > zswap/zram have been doing? It feels like an independent issue which
> > should be done separately?
>
> If you have 3 containers using compression space, and two of them have
> writeback enabled to a shared swapfile, the memory.swap.* controls
> need to work to manage fair access to that swapfile. They do not work
> if compression space itself is conflated in.
>
> Right now zswap entries actually consume physical swapfile space, even
> before writeback. Charging the space is correct. But the whole point
> is to decouple compression space from physical swap space.
>
> This is not something that can be done later. It would be a dramatic
> user-visible change to how the resource is categorized and managed.
>
> > > 2. You must make the compression space large enough to be outside the
> > > range where users can hit space limits before hitting memory limits.
> >
> > We may need a way to define 'large enough' at first.
>
> I've tried to lay this out in the other thread, and highlighted the
> usability issues that result from hitting compression space limits
> prematurely. It's kind of your call whether you want to seriously
> engage with this or not.
>
> But ultimately it's your claim that a static size can be made to work,
> so it's on you to make a convincing case.
>
> > > That also means not allowing setups where this is possible.
> >
> > And the limit is only an optional knob. If the admin does not set it,
> > the device grows to the full address space, so there is no space limit
> > to hit at all. It already behaves the way you want by default. The knob
> > is only for admins who want a ceiling, they can use it or not. I hope
> > this would not be a problem for your use case.
>
> No, I've laid this out already as well.
>
> This isn't about "my" usecase. It's about designing a coherent
> interface that works well with a large number of usecases, and other
> pieces of kernel infrastructure commonly used in conjunction.
>
> The other proposal in the room needs no such interface. The burden of
> proof for adding one is on you.
>
> > > [1] https://lore.kernel.org/linux-mm/aqLi6cIjD2wJwk0B@xxxxxxxxxxx/