Re: [PATCH v5 04/11] mm: zswap: add range lookup for large-folio swapin

From: Yosry Ahmed

Date: Thu Jul 23 2026 - 12:58:28 EST


> >> @@ -1595,13 +1611,19 @@ int zswap_load(struct folio *folio)
> >> return -ENOENT;
> >>
> >> /*
> >> - * Large folios should not be swapped in while zswap is being used, as
> >> - * they are not properly handled. Zswap does not properly load large
> >> - * folios, and a large folio may only be partially in zswap.
> >> + * A large folio reaches zswap_load() only when its whole range is
> >> + * expected to be on disk: PMD swap-entry consumers split before
> >> + * calling into PMD-order swapin whenever any slot is still in zswap.
> >> + * Confirm the range is entirely absent from zswap and return -ENOENT
> >> + * so the caller reads it from disk; if a slot is unexpectedly still in
> >> + * zswap, fail the read rather than return partially-initialized data.
> >> */
> >> - if (WARN_ON_ONCE(folio_test_large(folio))) {
> >> - folio_unlock(folio);
> >> - return -EINVAL;
> >> + if (folio_test_large(folio)) {
> >> + if (zswap_is_present(swp, folio_nr_pages(folio))) {
> >
> > Is dropping the warning here intentional (for the folio_test_large() &&
> > zswap_is_present() case)?
> >
>
>
> Yes, so we can end up in a race, which should be handled gracefully.
>
> For example, lets say we have zswap writeback enabled, which means we
> can end up in a state where we have a PMD swap entry and 511 of the 512
> slots have been written to disk, but 1 slot (slot X) is still in zswap.
>
> We can then have the following race:
>
>
> CPU A: PMD swap-in CPU B: split-PTE swap-in
> ------------------ ------------------------
>
> Checks swap cache: empty
>
> Faults slot X
> Adds order-0 folio F
> zswap_load(F):
> removes X from zswap
> marks F dirty
>
> Checks zswap range: empty
> A is preempted
>
> Unmaps/reclaims F
> zswap_store(F):
> puts X back in zswap
> Removes F from swap cache
>
> A resumes
> Allocates large swap-cache folio G
> zswap_load(G) finds X in zswap

Why don't we check the zswap range after allocating a folio in the
swap cache? I am assuming at this point we have the folio locked and
the result should be stable?

>
> The correct action would be to reject the PMD order read and fall back
> to per-page loading.
>
> I am bit torn about what to do for zswap here. The series is quite big
> already. Alexandre is looking at adding support for PMD swap, but will
> send patches once this series gets merged. My initial versions 1
> and 2, basically stopped installing PMD swap entry if zswap was ever enabled.
> From v3, I used Alexandre's suggestion to do zswap_is_present() test
> so that PMD swap entry can keep on working.
>
> Both of these paths are temporary till Alex sends his series. I
> do feel my initial version was simpler but will basically stop
> working if zswap is ever enabled. Do you have any suggestions on
> what your preference is for zswap?

Even with PMD zswap load support, we still have to deal with the case
where some order-0 pages are in zswap and some are on disk, right? I
don't see how that would eliminate the case you mentioned above.