Re: [PATCH v3 0/2] mm/swap: skip empty clusters in the swapoff scan
From: Baoquan He
Date: Fri Aug 07 2026 - 04:47:01 EST
On 08/07/26 at 03:41pm, Youngjun Park wrote:
> On Thu, Aug 06, 2026 at 01:06:55PM -0700, Andrew Morton wrote:
> > On Fri, 7 Aug 2026 04:32:26 +0900 Youngjun Park <youngjun.park@xxxxxxx> wrote:
> >
> > > find_next_to_unuse() walks a swap device one offset at a time. Slot
> > > state now lives in a per cluster swap table, so patch 2 dismisses an
> > > empty cluster with one counter read instead of SWAPFILE_CLUSTER table
> > > reads.
> >
> > Thanks.
> >
> > Can you help us understand how significant this change is for users?
> > If "not very" then I'd prefer to defer consideraton of the series until
> > after 7.3-rc1.
>
> Hello Andrew
>
> "Not very" in the common case, though there is a case where the win is clear.
> No bug and no user report.
>
> For now I would rather defer to after 7.3-rc1.
>
> And for your reference, here is the details.
>
> Every swapoff does a little less work now, because the scan steps over an
> unused area one cluster at a time.
> But IMHO most of the swapoff time goes to unuse_mm() and to reading the pages back in.
>
> The gain shows on a large swap device that is almost empty, when the last
> pages still in use are near the end of it. The scan has to walk up to
> them, and today it looks at every slot on the way. Now the empty clusters
> in between are skipped in one step.
>
> I have no measured times yet, since that case has to be set up on purpose.
> What I did is the arithmetic for the case that skips best,
> For example 1T of swap with 256M slots with SWAPFILE_CLUSTER = 512
> and everything free but the far end:
>
> - today: 256M table reads
> - with the skip: 512K counter reads
Maybe just use time to measure swapoff time consuming, just like below
as I did on a kvm guest, I guess a bare metal machine with larger system
ram could be more obvious?
root@fedora:~# free -h
total used free shared buff/cache available
Mem: 3.8Gi 181Mi 3.6Gi 924Ki 72Mi 3.7Gi
Swap: 2.0Gi 18Mi 2.0Gi
root@fedora:~# time swapoff /dev/vdb
real 0m0.101s
user 0m0.001s
sys 0m0.017s
root@fedora:~# swapon /dev/vdb
root@fedora:~# time swapoff /dev/vdb
real 0m0.014s
user 0m0.003s
sys 0m0.001s
Not sure if Andrew is asking for this.
>
> That should be around half a second of scan saved.
Yeah, a concrete number is shown.