Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload

From: Kairui Song

Date: Sat Sep 26 2026 - 06:16:48 EST


On Sat, Sep 26, 2026 at 12:55 PM Matthias Goergens
<matthias.goergens@xxxxxxxxx> wrote:
>
> This RFC adds offload-only swap areas for deliberate cold-page offload to
> backends unsuitable for pressure reclaim. ZFS zvol swap has documented
> deadlocks under memory pressure [3]; compressed swap is another target
> because writes may need memory despite free logical slots.
>
> Explicit proactive reclaim can use these areas alongside conventional
> swap, ordered by priority. Ordinary reclaim cannot initiate non-zero
> writes to them, including through retained swap entries. Conventional
> capacity is not reserved for emergencies. Operators choose the policy;
> the kernel does not measure headroom or make an allocating backend safe.
>
> TMO/Senpai [4] and DAMON_RECLAIM [5] are related cold-page reclaim
> approaches. This RFC admits offload-only swap through memory.reclaim and
> per-node reclaim, but not DAMON. Other related work includes per-cgroup
> zswap writeback control [6], Virtual Swap Space v4 [1] and swap tiers v10
> [2].
>
> The util-linux companion [8] proposes swapon --offload-only and an fstab
> option. Older swapon silently ignores the fstab token, so persistent
> activation needs discussion. Apply the separate i915 fix [9] first to avoid
> a pre-existing folio-lock leak when shmem writeback is skipped.
>
> The series fixes zram selftest device tracking and error reporting, adds
> the swap policy with DRM eligibility checks, and tests routing, workingset
> activation and retained-entry write refusal. It is based on mm-new at
> 995829088503.
>
> Feedback is particularly welcome on:
>
> - Activation-time eligibility versus backend placement or migration:
> is retained-entry refusal, with possible reclaim churn or OOM,
> acceptable without migration?
> - Eligible-capacity accounting, including overcommit limits and OOM
> scoring; cache recovery, cluster invalidation and workingset activation.
> - Persistent activation and visibility: current swap listings do not
> expose the policy.
>
> Validation includes x86 builds, affected arm64 MTE objects and builds
> without swap or memory cgroups. QEMU tests covered routing, retained-entry
> recovery, data integrity and cleanup, with negative controls for write
> refusal and teardown failure.
>
> I have also been using this policy on my own machine with experimental
> bcachefs swap support, with no problems observed so far. Deliberate stress
> testing is confined to VMs; I do not deliberately stress-test this machine.
>
> Changes since v1, incorporating Sashiko's public review [7]:
>
> - Improve selftest isolation, retained-swap measurement and memlock skips.
> - Track allocated zram devices, wait for udev probes and report cleanup
> errors without deleting data through a mount that could not be released.
> - Clarify activation, capacity reporting and conventional fallback limits.
>
> [1] https://patchew.org/linux/20260825153238.2695446-1-nphamcs@xxxxxxxxx/
> [2] https://patchew.org/linux/20260713025644.170839-1-youngjun.park@xxxxxxx/
> [3] https://openzfs.github.io/openzfs-docs/Project%20and%20Community/FAQ.html#using-a-zvol-for-a-swap-device-on-linux
> [4] https://engineering.fb.com/2022/06/20/data-infrastructure/transparent-memory-offloading-more-memory-at-a-fraction-of-the-cost-and-power/
> [5] https://www.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html
> [6] https://www.kernel.org/doc/html/latest/admin-guide/mm/zswap.html
> [7] https://sashiko.dev/#/patchset/20260914104501.3960616-1-matthias.goergens@xxxxxxxxx
> [8] https://github.com/util-linux/util-linux/pull/4633
> [9] https://lore.kernel.org/all/20260914105606.3997649-1-matthias.goergens@xxxxxxxxx/
>
> v1: https://lore.kernel.org/all/20260914104501.3960616-1-matthias.goergens@xxxxxxxxx/
>
> Matthias Goergens (4):
> selftests: zram: track owned devices and report cleanup failures
> mm: restrict offload-only swap to proactive reclaim
> selftests: zram: cover offload-only swap policy
> selftests: zram: cover retained offload-only entries
>
> Documentation/mm/swap.rst | 105 ++++
> drivers/gpu/drm/i915/gem/i915_gem_shrinker.c | 2 +-
> .../gpu/drm/i915/gem/selftests/huge_pages.c | 6 +-
> drivers/gpu/drm/msm/msm_gem_shrinker.c | 2 +-
> drivers/gpu/drm/panthor/panthor_gem.c | 2 +-
> drivers/gpu/drm/ttm/ttm_backup.c | 2 +-
> drivers/gpu/drm/xe/tests/xe_bo.c | 2 +-
> include/linux/swap.h | 29 +-
> include/linux/vm_event_item.h | 1 +
> mm/memcontrol.c | 20 +-
> mm/page_io.c | 24 +-
> mm/swapfile.c | 123 ++++-
> mm/vmscan.c | 31 +-
> mm/vmstat.c | 1 +
> tools/testing/selftests/zram/.gitignore | 2 +
> tools/testing/selftests/zram/Makefile | 4 +-
> tools/testing/selftests/zram/README | 20 +-
> tools/testing/selftests/zram/config | 10 +-
> tools/testing/selftests/zram/settings | 1 +
> tools/testing/selftests/zram/swap_offload.c | 425 +++++++++++++++++
> .../selftests/zram/workingset_offload.c | 207 ++++++++
> tools/testing/selftests/zram/zram.sh | 9 +
> tools/testing/selftests/zram/zram01.sh | 9 +-
> tools/testing/selftests/zram/zram02.sh | 9 +-
> tools/testing/selftests/zram/zram03.sh | 174 +++++++
> tools/testing/selftests/zram/zram04.sh | 164 +++++++
> tools/testing/selftests/zram/zram05.sh | 447 ++++++++++++++++++
> tools/testing/selftests/zram/zram_lib.sh | 189 ++++++--
> 28 files changed, 1926 insertions(+), 94 deletions(-)

Hi Matthias

That's a lot of changes for a rather limited usage. IIUC what it want
to archive is, for different kind of reclaim, swap operations should
only go to certain devices?

This really sounds like another usage of the swap tiering design: A
default per-cgroup tier setup, and a one time tier limit during
proactive reclaim.

Just like we already have a "swappiness=" parameter in the reclaim
interface, perhaps adding a "swap.tier=" would be better and much
cleaner to achieve the same goal? Based on Youngjun's work here:
https://lore.kernel.org/all/20260916183437.2946306-1-youngjun.park@xxxxxxx/