[PATCH 0/6] mm, swap: charge xswap for physical backing (xswap phase III)

From: Baoquan He

Date: Fri Oct 02 2026 - 21:11:39 EST


An xswap entry reserves no swap space: its page is in zswap, which costs
memory, not disk. It was still charged against memcg->swap at allocation,
so memory.swap.current counted pages that never reached a disk, and a
cgroup with memory.swap.max at 0 could not reclaim its anon memory at all.

The charge belongs where the backing is taken. Record the owner at
allocation; charge only when the entry takes a backend slot, and release
the charge with the slot.

Phase III of three, posted separately because the charging semantics are
still under discussion. This series makes memory.swap.current count only
the pages that hold a backend slot; it used to count every swap entry, and
an entry always held a slot until xswap separated the two.

Design notes
------------

Patch 1 splits __mem_cgroup_try_charge_swap() into get / charge / record /
uncharge / put, with no behaviour change, so the record and the charge can
happen at different times. The owner is a private ID, not a pinned memcg:
an ID is what the swap table holds, and the uncharge path resolves it back
under RCU. Patch 4 stops mem_cgroup_get_nr_swap_pages() from gating on the
physical margin for a cgroup that can zswap.

The patches
-----------

1 split the swap memcg charge helpers
2 do not charge zswap-backed xswap entries
3 charge an xswap entry when it gets physical backing
4 don't gate xswap on the physical swap free count
5 do not retake the cluster lock when uncharging an xswap slot
6 drop a refused xswap backend run directly

5 and 6 fix bugs in 3: a spinlock taken twice, and a refused run given back
through a reverse mapping that is not installed yet.

Testing
-------

qemu KVM guest, 8G RAM, xswap.max=10G, zram as the backend. memhog faults
<total_gb> of anon, fills it with a fixed pattern, and holds it. The
invariant through all of it:

memory.swap.current / 4096 == backend Used / 4

Both count pages holding a backend slot, so over-charge, under-charge and a
missed uncharge all show up here.

# echo 1 > /sys/module/zswap/parameters/enabled
# echo 100 > /sys/kernel/mm/xswap/create
# mkswap /dev/zram0; swapon -p 0 /dev/zram0
# mkdir -p /sys/fs/cgroup/xswap_limit
# echo 2G > .../memory.max; echo max > .../memory.swap.max
# echo max > .../memory.zswap.max
# ( echo $BASHPID > .../cgroup.procs
# exec env MEMHOG_FILL=pattern numactl --cpunodebind=0 ./memhog 4 600 ) &

With zswap on nothing reaches the backend and nothing is charged. With
zswap off the two sides agree (526429 == 526429). memory.swap.max is
enforced at the backend: at 0, a device-less cgroup is killed, but with an
xswap device the pages go to zswap, cost no swap space, and reclaim keeps
making progress. Destroying the device returns the charge and the backend
Used with it, workload alive, over ten cycles. Also exercised: the same
invariant on an ordinary swap device with no xswap device.

Changelog
=========
RFC -> v1:

- Taken out of the RFC as a series of its own. The direction for charging
is still under discussion, and the writeback series should not wait on
it.

Baoquan He (2):
mm, swap: do not retake the cluster lock when uncharging an xswap slot
mm, swap: drop a refused xswap backend run directly

Nhat Pham (4):
mm, swap: split the swap memcg charge helpers
mm, swap: do not charge zswap-backed xswap entries
mm, swap: charge an xswap entry when it gets physical backing
mm, swap: don't gate xswap on the physical swap free count

.../admin-guide/cgroup-v1/memcg_test.rst | 2 +-
include/linux/memcontrol.h | 6 +
include/linux/swap.h | 66 ++++++-
mm/memcontrol-v1.c | 3 +-
mm/memcontrol.c | 164 +++++++++++++-----
mm/swapfile.c | 138 ++++++++++++++-
6 files changed, 318 insertions(+), 61 deletions(-)


base-commit: 456d37694d1e05a1b9ddbd1e493bcd6cf2016b4f
--
2.54.0