Re: [PATCH v3 6/7] mm/mglru: move folios from oldest gen to second-oldest gen from head to tail
From: Barry Song
Date: Thu Sep 03 2026 - 22:29:40 EST
On Fri, Sep 4, 2026 at 9:53 AM Baolin Wang
<baolin.wang@xxxxxxxxxxxxxxxxx> wrote:
>
>
>
> On 9/2/26 7:24 AM, Barry Song (Xiaomi) wrote:
> > For reclamation, it makes sense to reclaim folios from tail to
> > head, as folios near the head are relatively hot. However, when
> > moving folios from the oldest generation to the second-oldest
> > generation, using the tail-to-head order would effectively cause
> > a cold/hot inversion.
> >
> > Signed-off-by: Barry Song (Xiaomi) <baohua@xxxxxxxxxx>
> > Reviewed-by: Baoquan He <baoquan.he@xxxxxxxxx>
> > Tested-by: Xueyuan Chen <xueyuan.chen21@xxxxxxxxx>
> > Reviewed-by: Lian Wang <lianux.mm@xxxxxxxxx>
> > ---
> > mm/vmscan.c | 23 +++++++++++++++++++++--
> > 1 file changed, 21 insertions(+), 2 deletions(-)
> >
> > diff --git a/mm/vmscan.c b/mm/vmscan.c
> > index 1907a946840d..76dfa9575852 100644
> > --- a/mm/vmscan.c
> > +++ b/mm/vmscan.c
> > @@ -192,11 +192,27 @@ static inline void prefetchw_prev_lru_folio(struct folio *folio,
> > prefetchw(&prev->flags);
> > }
> > }
> > +
> > +static inline void prefetchw_next_lru_folio(struct folio *folio,
> > + struct list_head *base)
> > +{
> > + if (folio->lru.next != base) {
> > + struct folio *next;
> > +
> > + next = list_entry(folio->lru.next, struct folio, lru);
> > + prefetchw(&next->flags);
> > + }
> > +}
>
> You did not mention in the commit message why this prefetch was added.
> Just curious, does it really help performance?
Hi Baolin,
Thanks for your review.
I had this discussion with Kairui and thought it might be
useful to keep it here:
https://lore.kernel.org/linux-mm/CAMgjq7AJnWGGH=FNuV7ynXjbffktsDRgdcyagLUJUwJf1ukaJQ@xxxxxxxxxxxxxx/
It might be arch-dependent. Some architectures could benefit
significantly from prefetching, while others might see little to
no impact.
for my x86 test, it has very slight improvement:
***** no-prefetch:
agetest:
...
gen 100: 2.421 ms
gen 101: 2.424 ms
gen 102: 2.420 ms
gen 103: 2.413 ms
Total: 248.418 ms
Average: 2.460 ms
agetest:
...
gen 100: 2.393 ms
gen 101: 2.396 ms
gen 102: 2.395 ms
gen 103: 2.392 ms
Total: 245.627 ms
Average: 2.432 ms
agetest:
...
gen 100: 2.433 ms
gen 101: 2.427 ms
gen 102: 2.432 ms
gen 103: 2.450 ms
Total: 249.186 ms
Average: 2.467 ms
**** has-prefetch:
agetest:
....
gen 100: 2.314 ms
gen 101: 2.310 ms
gen 102: 2.321 ms
gen 103: 2.303 ms
Total: 237.619 ms
Average: 2.353 ms
agetest:
gen 100: 2.342 ms
gen 101: 2.343 ms
gen 102: 2.339 ms
gen 103: 2.335 ms
Total: 239.929 ms
Average: 2.376 ms
agetest:
gen 100: 2.344 ms
gen 101: 2.347 ms
gen 102: 2.352 ms
gen 103: 2.348 ms
Total: 241.188 ms
Average: 2.388 ms
Basically, it’s 2.3xx vs. 2.4xx, lower is better.
Best Regards
Barry