[PATCH v1 0/2] mm: fix workingset refaults in the zswap writeback path

From: Alexandre Ghiti

Date: Mon Aug 17 2026 - 11:55:14 EST


When an anonymous folio is reclaimed, workingset_eviction() stores a
"shadow" (the eviction cookie) in the swap slot so that a later swap-in
can be recognised as a refault and, if the refault distance is short
enough, the page can be re-activated. This is how anon workingset/refault
detection has worked since commit aae466b0052e ("mm/swap: implement
workingset detection for anonymous LRU").

zswap writeback breaks this in two independent ways:

- Over-count at writeback: the shrinker allocates a buffer folio in the
swap cache, and the allocation path counts that folio as a refault.

- Lost eviction cookie at reclaim: adding the buffer to the swap cache
overwrites the slot's shadow, so the original cookie is lost; when the
buffer folio is finally reclaimed a fresh, inaccurate cookie is minted
in its place.

This series preserves the shadow within zswap itself, without any new
swap-table or swap-slot state. At writeback, instead of erasing the freed
zswap entry, the captured shadow is parked in the zswap tree in its place,
so it outlives the writeback buffer folio.

Finally, the refault evaluation is moved out of the swap-cache allocator
into the swap-in callers, so allocating the writeback buffer is no longer
miscounted as a refault.

Results
-------

Measured with a sysbench OLTP (MariaDB) workload in a memory cgroup sized
so the dataset and InnoDB buffer pool both overcommit it, with the zswap
shrinker on so entries are continuously written back to an NVMe swap
device (classic LRU; MGLRU off). Anon workingset counters over the
measured window, baseline vs this series, mean +/- stddev over 10 runs:

workingset_refault_anon 383,242 +/- 59,345 -> 199,355 +/- 27,042 -48%
workingset_activate_anon 54,890 +/- 11,742 -> 22,607 +/- 3,118 -59%
workingset_restore_anon 16,212 +/- 4,590 -> 7,599 +/- 1,190 -53%

Writeback volume is comparable (zswpwb 183k +/- 16k -> 178k +/- 15k), so
the reduction is not from doing less work. Normalised per transaction the
reduction holds (-53%/-49%/-39%) while the swap work per transaction is
unchanged. A kernel build under the same pressure moves all three
counters in the same direction.

The run-to-run variance of these counters drops as well.

Throughput is unaffected: over the same 10 runs, transactions/s is
18.50 +/- 1.08 -> 18.65 +/- 0.94, i.e. +0.8% with a 95% confidence
interval of +/- 7.7%.

Alexandre Ghiti (2):
mm/swap: refault on swap-in, not in the swap cache allocator
mm/zswap: preserve the workingset shadow across writeback

include/linux/zswap.h | 12 +++++++
mm/memory.c | 5 +++
mm/shmem.c | 5 +++
mm/swap_state.c | 30 ++++++++++++++--
mm/vmscan.c | 4 ++-
mm/workingset.c | 5 ++-
mm/zswap.c | 80 ++++++++++++++++++++++++++++++++++++++++++-
7 files changed, 136 insertions(+), 5 deletions(-)


base-commit: 626acb37cd445144f321f1b64cac9a93760fa716
--
2.53.0-Meta