Re: [RFC PATCH v2 01/10] mm: xswap support for zswap
From: Baoquan He
Date: Fri Aug 07 2026 - 05:11:35 EST
Hi Johannes,
On 08/05/26 at 10:17am, Johannes Weiner wrote:
> On Wed, Aug 05, 2026 at 03:53:24PM +0800, Baoquan He wrote:
> > From: Chris Li <chrisl@xxxxxxxxxx>
> >
> > Introduce extendable (virtual) swap device support ??? xswap.
> >
> > The current zswap requires a backing swapfile. The swap slot used
> > by zswap is not able to be used by the swapfile, wasting swapfile
> > space.
> >
> > An xswap device is a swapfile that only contains the swap header,
> > with the header indicating the size of the virtual swap space. There
> > is no swap data section, therefore no waste of swapfile space. Any
> > write to an xswap device will fail. To prevent accidental read or
> > write, bdev of swap_info_struct is set to NULL. Xswap devices set
> > the SSD flag because there is no rotational disk access when using
> > zswap.
> >
> > Zswap writeback is disabled if all swapfiles in the system are
> > xswap devices (tracked via nr_real_swapfiles).
> >
> > How to create an xswap device:
> > touch swap.1G
> > truncate -s 1G swap.1G
> > mkswap swap.1G
> > dd if=swap.1G of=xswap.1G bs=4096 count=1
> > # xswap.1G is 4K on disk but reports 1G capacity
> > swapon xswap.1G
>
> Sigh.
>
> Why does the user have to go through this dance?
>
> Why does the user have to decide in advance what size the space needs
> to be?
>
> You point out no inherent limit to how much can be compressed, so
> there is no reason to make userspace decide on an arbitrary one.
>
> There is no reason to tie an address space that can be managed
> transparently inside the kernel to TWO named files on disk.
Thanks for looking into this.
The file-based creation dance is there only because this is RFC —
I wanted to reuse the existing swapon path so the core grow/shrink
machinery could be measured and tested without also designing a new
userspace interface. I agree it's not the right final interface.
The direction I'm thinking for the next revision:
- Drop the file requirement entirely. An xswap device has no backing
store, so there is no reason it needs a file.
- Use totalram_pages as the initial per-device size. Chris suggested
this, and it's a natural bound: if all anonymous memory is swapped
out, that is the maximum number of swap entries zswap will ever need,
assuming a reasonable compression ratio. The hard upper limit could
be 2 times of system RAM, or the max system RAM memory hotplug can
add to.
Doing this because we need consider swap.tier support. A single global
xswap device in swap.tier would mean all memcgs compress into the
same device — there is only one swap entry namespace. With per-device
xswap instances, swap.tier can bind different memcgs to different xswap
devices, giving each its own swap slot namespace. Total isolation on slot
usage, no cross-memcg interference.
-----
Hi Chris, Joungjun,
Please correct me if I misunderstood the swap.tier concept and xswap
use case in there.)
-----
For creation, something like:
1.
echo $((4 * 1024 * 1024 * 1024)) > /sys/kernel/mm/xswap/create
or
2.
even simpler, with automatic sizing:
echo 1 > /sys/kernel/mm/xswap/enable
(use totalram_pages as the default size.)
3.
swapon -t xswap xswap0
(use totalram_pages as the default size.)
I'm open to other ideas. If anyone have a preference for the interface,
I'd like to hear it.
>
> There are plenty of past discussions on this very topic. I don't see
> the point in resubmitting the same thing under different names,
> without even a reference to previous discussions.
>
> As a side note: if you have to add "(virtual)" after every instance of
> "extendable", then maybe "extendable" is a terrible name and you
> should just call it "virtual".
"extendable (virtual)" wasn't meant to explain one with the other.
Chris prefers "extendable", you prefer "virtual" — I put both in the
cover letter so the community could weigh in. I don't have a strong
preference to either.
I'll address all of this in RFC v3 — drop the file requirement,
auto-size to totalram_pages (with grow/shrink for dynamic adjustment
on top), and settle the name. The goal of RFC v2 was to get the
core mechanics reviewed; I think that part is in reasonable shape,
and the interface is exactly the kind of thing I was hoping to get
feedback on. Thanks for providing it.
Thanks
Baoquan