Re: [PATCH v3 6/7] mm/mglru: move folios from oldest gen to second-oldest gen from head to tail
From: Baolin Wang
Date: Thu Sep 03 2026 - 22:49:52 EST
On 9/4/26 10:28 AM, Barry Song wrote:
On Fri, Sep 4, 2026 at 9:53 AM Baolin Wang
<baolin.wang@xxxxxxxxxxxxxxxxx> wrote:
On 9/2/26 7:24 AM, Barry Song (Xiaomi) wrote:
For reclamation, it makes sense to reclaim folios from tail to
head, as folios near the head are relatively hot. However, when
moving folios from the oldest generation to the second-oldest
generation, using the tail-to-head order would effectively cause
a cold/hot inversion.
Signed-off-by: Barry Song (Xiaomi) <baohua@xxxxxxxxxx>
Reviewed-by: Baoquan He <baoquan.he@xxxxxxxxx>
Tested-by: Xueyuan Chen <xueyuan.chen21@xxxxxxxxx>
Reviewed-by: Lian Wang <lianux.mm@xxxxxxxxx>
---
mm/vmscan.c | 23 +++++++++++++++++++++--
1 file changed, 21 insertions(+), 2 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 1907a946840d..76dfa9575852 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -192,11 +192,27 @@ static inline void prefetchw_prev_lru_folio(struct folio *folio,
prefetchw(&prev->flags);
}
}
+
+static inline void prefetchw_next_lru_folio(struct folio *folio,
+ struct list_head *base)
+{
+ if (folio->lru.next != base) {
+ struct folio *next;
+
+ next = list_entry(folio->lru.next, struct folio, lru);
+ prefetchw(&next->flags);
+ }
+}
You did not mention in the commit message why this prefetch was added.
Just curious, does it really help performance?
Hi Baolin,
Thanks for your review.
I had this discussion with Kairui and thought it might be
useful to keep it here:
https://lore.kernel.org/linux-mm/CAMgjq7AJnWGGH=FNuV7ynXjbffktsDRgdcyagLUJUwJf1ukaJQ@xxxxxxxxxxxxxx/
It might be arch-dependent. Some architectures could benefit
significantly from prefetching, while others might see little to
no impact.
for my x86 test, it has very slight improvement:
***** no-prefetch:
agetest:
...
gen 100: 2.421 ms
gen 101: 2.424 ms
gen 102: 2.420 ms
gen 103: 2.413 ms
Total: 248.418 ms
Average: 2.460 ms
agetest:
...
gen 100: 2.393 ms
gen 101: 2.396 ms
gen 102: 2.395 ms
gen 103: 2.392 ms
Total: 245.627 ms
Average: 2.432 ms
agetest:
...
gen 100: 2.433 ms
gen 101: 2.427 ms
gen 102: 2.432 ms
gen 103: 2.450 ms
Total: 249.186 ms
Average: 2.467 ms
**** has-prefetch:
agetest:
....
gen 100: 2.314 ms
gen 101: 2.310 ms
gen 102: 2.321 ms
gen 103: 2.303 ms
Total: 237.619 ms
Average: 2.353 ms
agetest:
gen 100: 2.342 ms
gen 101: 2.343 ms
gen 102: 2.339 ms
gen 103: 2.335 ms
Total: 239.929 ms
Average: 2.376 ms
agetest:
gen 100: 2.344 ms
gen 101: 2.347 ms
gen 102: 2.352 ms
gen 103: 2.348 ms
Total: 241.188 ms
Average: 2.388 ms
Basically, it’s 2.3xx vs. 2.4xx, lower is better.
Could you add this info to the commit message? (IIUC, someone tried to remove the prefetch in MM, since they found that prefetch doesn't seem to help much on modern CPUs.)