Re: [PATCH v3] mm: bypass swap readahead for zswap
From: Yosry Ahmed
Date: Wed Oct 07 2026 - 17:37:27 EST
On Wed, Oct 7, 2026 at 8:39 AM Alexandre Ghiti <alex@xxxxxxxx> wrote:
>
> Commit 0bcac06f27d7 ("mm, swap: skip swapcache for swapin of synchronous
> device") made SWP_SYNCHRONOUS_IO devices (e.g. zram) skip swap readahead.
>
> zswap is the same kind of in-memory, synchronous backend as zram, not a
> swap device flagged SWP_SYNCHRONOUS_IO so it still goes through
> swapin_readahead().
>
> Here are the results from bypassing readahead for zswap too: it was
> measured with a kernel build (make -j16) in a memcg, zswap=zstd, shrinker
> off, on Sapphire Rapids and 3 iterations.
>
> 768M memcg (sustained swap thrash):
> metric mm-new + bypass delta
> build time (s) 405.0 341.7 -15.6%
> zswap-in (GB) 79.5 53.0 -33%
> zswap-out (GB) 144.8 115.6 -20%
> swap readahead (pages) 6.79M 0.45M -93%
> swap_ra hit (%) 72.1 89.9 +18pp
>
> 1G memcg (light pressure, build not memory-bound):
> metric mm-new + bypass delta
> build time (s) 177.7 176.0 ~same (no regression)
> zswap-in (GB) 10.2 7.5 -26%
> zswap-out (GB) 27.7 25.1 -9%
> swap readahead (pages) 1.07M 0.08M -93%
> swap_ra hit (%) 68.6 87.2 +19pp
>
> Similar gains were observed on an AMD EPYC 7D13 host.
>
> The gain is from no longer prefetching pages that are pointless for an
> in-memory backend: readahead inflates anon residency and thrashes the
> page cache (file pages get evicted and re-read), lengthens each fault by
> synchronously (de)compressing a cluster of neighbours, and adds
> compression traffic when those extra pages are reclaimed.
>
> Bypassing swap readahead for zswap therefore makes sense.
No objection here, I am just thinking about Kairui's LPC presentation
about how compressed swap can (in theory at least) still benefit from
readahead: https://lpc.events/event/20/contributions/2421/attachments/2116/4689/Better%20Anon%20(Swap)%20Readahead%20(2).pdf.
Seems like a big portion of the fault cost is still outside of
decompression and can benefit from readahead, but that isn't the case
today. I assume any future improvements to readahead and re-enabling
for zram would also cover zswap?
>
> Suggested-by: Usama Arif <usama.arif@xxxxxxxxx>
> Signed-off-by: Alexandre Ghiti <alex@xxxxxxxx>
> ---
> Changes in v3:
> - Go back to the RFC's single patch: readahead windows that mix zswap
> and disk entries will be dealt with later, the priority is to merge the
> zswap readahead bypass
> - Move the check into a swap helper, swap_entry_synchronous(), instead
> of open-coding it in do_swap_page() (David)
> - Bail out early when zswap was never enabled or the tree is empty, so
> disk-only systems pay a static branch (Kairui, Nhat)
> - Reuse zswap_is_present() rather than adding zswap_present_test()
> (Yosry)
> - Rebase on mm-new
>
> Changes in v2:
> - Keep readahead, but overlap the decompression of zswap neighbours with
> the disk I/O of the window, and bypass it when the target needs no
> disk I/O (Yosry, Nhat, Barry)
>
> v2: https://lore.kernel.org/all/20260722163536.773166-1-alex@xxxxxxxx/
> RFC: https://lore.kernel.org/all/20260624075700.751467-1-alex@xxxxxxxx/
>
> include/linux/zswap.h | 6 ++++++
> mm/memory.c | 4 ++--
> mm/swap.h | 25 +++++++++++++++++++++++++
> mm/zswap.c | 5 ++++-
> 4 files changed, 37 insertions(+), 3 deletions(-)
>
> diff --git a/include/linux/zswap.h b/include/linux/zswap.h
> index df6cafbe95dc..94746fb71bb6 100644
> --- a/include/linux/zswap.h
> +++ b/include/linux/zswap.h
> @@ -35,6 +35,7 @@ void zswap_lruvec_state_init(struct lruvec *lruvec);
> void zswap_folio_swapin(struct folio *folio);
> bool zswap_is_enabled(void);
> bool zswap_never_enabled(void);
> +bool zswap_is_present(swp_entry_t entry, unsigned int nr);
> #else
>
> struct zswap_lruvec_state {};
> @@ -73,6 +74,11 @@ static inline bool zswap_never_enabled(void)
> return true;
> }
>
> +static inline bool zswap_is_present(swp_entry_t entry, unsigned int nr)
> +{
> + return false;
> +}
> +
> #endif
>
> #endif /* _LINUX_ZSWAP_H */
> diff --git a/mm/memory.c b/mm/memory.c
> index 330cde31bf8b..396bc71a265f 100644
> --- a/mm/memory.c
> +++ b/mm/memory.c
> @@ -5030,8 +5030,8 @@ vm_fault_t do_swap_page(struct vm_fault *vmf)
> if (folio)
> swap_update_readahead(folio, vma, vmf->address);
> if (!folio) {
> - /* Swapin bypasses readahead for SWP_SYNCHRONOUS_IO devices */
> - if (data_race(si->flags & SWP_SYNCHRONOUS_IO))
> + /* Swapin bypasses readahead for SWP_SYNCHRONOUS_IO devices and zswap */
> + if (swap_entry_synchronous(si, entry))
> folio = swapin_sync(entry, GFP_HIGHUSER_MOVABLE,
> thp_swapin_suitable_orders(vmf) | BIT(0),
> vmf, NULL, 0);
> diff --git a/mm/swap.h b/mm/swap.h
> index d5bf21f517dc..b8181bd056ee 100644
> --- a/mm/swap.h
> +++ b/mm/swap.h
> @@ -94,6 +94,7 @@ static inline int mem_cgroup_swappiness(const struct mem_cgroup *memcg)
>
> #ifdef CONFIG_SWAP
> #include <linux/swapops.h> /* for swp_offset */
> +#include <linux/zswap.h> /* for zswap_is_present */
> #include <linux/blk_types.h> /* for bio_end_io_t */
>
> static inline unsigned int swp_cluster_offset(swp_entry_t entry)
> @@ -336,6 +337,24 @@ struct folio *swapin_sync(swp_entry_t entry, gfp_t flag, unsigned long orders,
> void swap_update_readahead(struct folio *folio, struct vm_area_struct *vma,
> unsigned long addr);
>
> +/**
> + * swap_entry_synchronous - check if @entry is read synchronously
> + * @si: the swap device of @entry
> + * @entry: the swap entry
> + *
> + * SWP_SYNCHRONOUS_IO devices and zswap complete a read in the faulting
> + * context, so swap readahead has no asynchronous I/O to overlap.
> + *
> + * Context: The caller must hold a reference on @si. The answer is a hint: the
> + * entry can be written back from zswap or freed concurrently.
> + */
> +static inline bool swap_entry_synchronous(struct swap_info_struct *si,
> + swp_entry_t entry)
> +{
> + return data_race(si->flags & SWP_SYNCHRONOUS_IO) ||
> + zswap_is_present(entry, 1);
> +}
> +
> #else /* CONFIG_SWAP */
>
> static inline struct swap_cluster_info *swap_cluster_get_and_lock(
> @@ -423,6 +442,12 @@ static inline void swap_update_readahead(struct folio *folio,
> {
> }
>
> +static inline bool swap_entry_synchronous(struct swap_info_struct *si,
> + swp_entry_t entry)
> +{
> + return false;
> +}
> +
> static inline int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
> {
> return 0;
> diff --git a/mm/zswap.c b/mm/zswap.c
> index eb3494491410..1dcf77d5edae 100644
> --- a/mm/zswap.c
> +++ b/mm/zswap.c
> @@ -1605,12 +1605,15 @@ bool zswap_store(struct folio *folio)
> * change under it.
> * Return: true if at least one slot in the range is in zswap.
> */
> -static bool zswap_is_present(swp_entry_t entry, unsigned int nr)
> +bool zswap_is_present(swp_entry_t entry, unsigned int nr)
> {
> pgoff_t offset = swp_offset(entry);
> struct xarray *tree = swap_zswap_tree(entry);
> unsigned long index = offset;
>
> + if (zswap_never_enabled() || xa_empty(tree))
> + return false;
> +
> /*
> * A pinned range is at most SWAPFILE_CLUSTER slots and is aligned to
> * its own size, so one tree covers all of it and a single lookup is
> --
> 2.53.0-Meta
>
>