Re: [PATCH v3 00/14] mm, swap: extendable swap devices (xswap)
From: Baoquan He
Date: Sun Sep 27 2026 - 21:36:38 EST
On 09/24/26 at 12:00pm, Klara Modin wrote:
> On 2026-09-16 18:19:07 +0800, Baoquan He wrote:
> > xswap is a swap device with no backing storage. Swapped-out pages live
> > in zswap. Its cluster_info[] array lives in a VM_SPARSE vmalloc area,
> > and the area is grown and shrunk on demand as swap usage changes.
> >
> > The problem being solved is the static size of compressed swap. Both
> > zram and zswap need the size fixed in advance, and neither gives memory
> > back when the workload shrinks. The solution should be a device whose
> > size can scale up/down as per usage. xswap does that by mapping the
> > metadata lazily instead of reserving it for the whole range.
> >
> > Design
> > ------
> > - si->cluster_info[] stays a plain array. Access is still
> > &si->cluster_info[offset / SWAPFILE_CLUSTER]: no per-access branch, no
> > RCU discipline, no tear-down state machine, no NULL return.
> > - Only an initial chunk is mapped at creation. The rest of the address
> > space is reserved, not allocated, so an idle device costs nothing.
> > - Growth is driven by allocation. When no free cluster is left and the
> > address space has room, the next chunk is mapped and added to the free
> > list. No userspace involvement.
> > - Shrink is driven by frees. The free tail is scanned, and whole chunks
> > are unmapped once the mapped range is at most half in use and several
> > chunks can go. One chunk is left mapped as slack, so the next
> > allocation does not map it straight back. A ceiling lowered below the
> > mapped range skips the half-in-use rule and is enforced at once.
> >
>
> > Size
> > ----
> > A device starts at 1xRAM, rounded down to the cluster. That costs
> > nothing, because the mapping is lazy. The underlying address space is
> > 2xRAM. An optional per-device cap,
> > /sys/kernel/mm/xswap/type<N>/limit, lets an admin lower the ceiling;
> > the excess is unmapped right away. Grow and shrink both work without
> > it. Creating a device requires zswap.
>
> So I can't set an xswap device to more than twice the RAM? I suppose I
> could create multiple xswap devices, but it would get tedious fast on
> systems which have a different amount of memory. Is there a particular
> reason for this limit? I think I could create an arbitrarily large xswap
> device with your previous version which needed the specially crafted
> swapfile (with only the header).
The 2xRAM setting was derived based on the current system condition.
Assume we take zstd which has the highest compression ration, the
compressed memory accounts for about 30%. So I set a max value 2xRAM
based on my own limited knowledge. I will fix that in v4, let user
decide.
And yes, in earlier verison, Jonhannes disliked the ghost swap file, so
I take a file-less sysfs interface way instead.
>
> As I wrote in the other thread, I would rather not have to set a limit
> at all, or at least have a limit I'm sure I won't reach.
I got it, it will be changed in v4.
Thanks a lot for your careful reviewing and testing.
Thanks
Baoquan