Re: [PATCH v3 0/2] mm/swap: skip empty clusters in the swapoff scan

From: Youngjun Park

Date: Fri Aug 07 2026 - 03:02:43 EST


On Fri, Aug 07, 2026 at 03:41:22PM +0900, Youngjun Park wrote:
> On Thu, Aug 06, 2026 at 01:06:55PM -0700, Andrew Morton wrote:
> > On Fri, 7 Aug 2026 04:32:26 +0900 Youngjun Park <youngjun.park@xxxxxxx> wrote:
> >
> > > find_next_to_unuse() walks a swap device one offset at a time. Slot
> > > state now lives in a per cluster swap table, so patch 2 dismisses an
> > > empty cluster with one counter read instead of SWAPFILE_CLUSTER table
> > > reads.
> >
> > Thanks.
> >
> > Can you help us understand how significant this change is for users?
> > If "not very" then I'd prefer to defer consideraton of the series until
> > after 7.3-rc1.
>
> Hello Andrew
>
> "Not very" in the common case, though there is a case where the win is clear.
> No bug and no user report.

Something I forgot to mention,

This only affects swapoff.
and few users run swapoff often, so the impact is limited either way.

> For now I would rather defer to after 7.3-rc1.
>
> And for your reference, here is the details.
>
> Every swapoff does a little less work now, because the scan steps over an
> unused area one cluster at a time.
> But IMHO most of the swapoff time goes to unuse_mm() and to reading the pages back in.
>
> The gain shows on a large swap device that is almost empty, when the last
> pages still in use are near the end of it. The scan has to walk up to
> them, and today it looks at every slot on the way. Now the empty clusters
> in between are skipped in one step.
>
> I have no measured times yet, since that case has to be set up on purpose.
> What I did is the arithmetic for the case that skips best,
> For example 1T of swap with 256M slots with SWAPFILE_CLUSTER = 512
> and everything free but the far end:
>
> - today: 256M table reads
> - with the skip: 512K counter reads
>
> That should be around half a second of scan saved.

+ benefit.

Reclaim can put a cached folio back into a page table with the slot it
already had, so try_to_unuse() retries and the scan starts over. The
saving then applies once per pass.