Re: [PATCH RFC v2 00/15] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup

From: Kairui Song

Date: Mon Sep 28 2026 - 05:40:17 EST


On Mon, Sep 28, 2026 at 5:12 PM Xiang Liu <liuxiang.277@xxxxxxxxxxxxx> wrote:
>
> Hi Kairui,
>
> I tested the complete MGLRU-FG RFC v2 series on an x86_64 machine with
> a SATA SSD. I found that the LevelDB Scan/Get result changes with the
> relative intensity of GET and SCAN.
>
> Setup: base commit 9958290885035431e2494159f555e332c3c14906, about
> 28 GiB LevelDB, 2.5 GiB memcg, swap off, Zipf 0.99, SCAN length 10000,
> 240 seconds warm-up plus 240 seconds measurement. Each untraced ABBA
> sample used a fresh database copy, page-cache drop, reboot and cooldown.
>
> Untraced ABBA results
> =====================
>
> Workload GET/SCAN Baseline FG Delta
> 1 SCAN + 3 GET 92.5 4688.62 4425.82 -5.61%
> 4 mixed workers, 99.5/0.5 195.2 6200.60 6385.57 +2.98%
> 1 SCAN + 7 GET 237.2 9574.97 9831.70 +2.68%
>
> The first workload is the original four-thread layout: one continuous
> SCAN thread and three GET threads. The second is a custom workload in
> which every worker chooses 99.5% GET or 0.5% SCAN per request. The third
> keeps one continuous SCAN thread and increases GET threads to seven.
>
> In the original workload, FG changed GET throughput by -5.74% and SCAN
> throughput by +6.91%. The two ABBA position pairs were -6.21% and
> -5.00%. A separate 5 GiB memcg repeat was also negative: -9.07% and
> -7.67%.
>
> Traced results
> ==============
>
> I ran one fresh-boot baseline/FG pair for each workload. To limit probe
> overhead, I selected about 1/1024 of file offsets using a stable hash of
> inode and page index.
>
> Initial placement of newly loaded GET folios
> --------------------------------------------
>
> We tracked new file pages loaded by GET requests and recorded where
> MGLRU initially placed each page: the oldest generation or the
> second-oldest generation.
>
> Workload / kernel Sampled Oldest Second-oldest
> 1S+3G / baseline 1452 186 (12.8%) 1266 (87.2%)
> 1S+3G / FG 1371 1052 (76.7%) 319 (23.3%)
> Mixed / baseline 1782 199 (11.2%) 1583 (88.8%)
> Mixed / FG 1662 1312 (78.9%) 350 (21.1%)
> 1S+7G / baseline 2719 469 (17.2%) 2250 (82.8%)
> 1S+7G / FG 2677 2107 (78.7%) 570 (21.3%)
>
> For example, the original 1-SCAN/3-GET baseline run sampled 1452 newly
> loaded GET folios. Of these, 1266 entered the second-oldest generation
> and 186 entered the oldest. With FG, 1052 of 1371 sampled folios entered
> the oldest generation.
>
> The same change appears in all three workloads: mainline places most
> new GET folios in the second-oldest generation, whereas FG places most
> of them in the oldest generation.
>
> FG promotion was active in all three workloads:
>
> Workload Later GET Moved to the promoted generation
> changed gen during reclaim scan
> 1 SCAN + 3 GET 38 548
> 4 mixed workers 32 558
> 1 SCAN + 7 GET 70 936
>
> However, FG also removed more sampled GET folios before another GET
> reused the same in-memory folio:
>
> Removed after one GET Later reloaded
> 1 SCAN + 3 GET 81.5% -> 88.4% 44.7% -> 48.6%
> 4 mixed workers 73.4% -> 90.2% 48.3% -> 52.7%
> 1 SCAN + 7 GET 84.5% -> 88.2% 56.2% -> 61.1%
>
> The arrows show baseline -> FG. "Later reloaded" means a later GET
> demanded the same file offset after it had been removed from page cache.
>
> The request-level counters explain why the final results differ:
>
> FG change per request GET cache miss GET I/O wait SCAN cache miss
> 1 SCAN + 3 GET +2.29% +6.87% -15.97%
> 4 mixed workers -5.68% -4.19% -6.93%
> 1 SCAN + 7 GET -2.89% -0.32% -8.26%
>
> FG improved the SCAN side in every workload. In the original workload,
> the GET miss and I/O increases were larger than the SCAN gain, producing
> the 5.61% total regression. When GET was denser relative to SCAN, GET
> misses also decreased and FG became a net win.
>
> My interpretation is that FG introduces a cold-start window for a new
> GET folio:
>
> first GET -> oldest generation
> -> a later GET promotes it, if it is still present
> -> otherwise reclaim removes it first
>
> The outcome depends on whether GET reuse arrives before the continuous
> SCAN stream consumes the oldest generation. The trace shows that both
> promotion and early removal occur in practice.
>
> This machine uses SATA, while the positive results in the cover letter
> used NVMe. Storage latency may also change the timing among page-cache
> insertion, LRU drain, the next GET and reclaim.
>
> Is this cold-start window an expected trade-off of the current design?
> Would limited initial protection for refs=1 folios, or a decision based
> on observed refault/reuse distance, be worth testing?

Hi Xiang, thanks for testing:

Well I mentioned this in the cover letter:

===

Interestingly, mainline MGLRU already beats CLRU on this one after a
recent change in lru_gen_folio_seq that bumps new folios with refs == 1
to the second-oldest generation. That change accidentally gave random
reads a higher hotness level while making sequential reads much colder:
sequential reads involve many readahead hits, and readahead folios start
with refs == 0, so they're already in the oldest generation, and
folio_mark_accessed() on a readahead hit has almost no effect on generations
in mainline MGLRU. Meanwhile, all direct-hit (random read) folios start
with refs == 1 in the second-oldest generation. As a result, the scan-get
test natively protects the "get" part and sacrifices the "scan" part.
That's not the best solution though. It's unreliable because it depends
on LRU drain timing, and it hurts workloads where the sequential part is
actually hotter (any workload involving a hotter large file and many
small cold files will be affected).

===

That actually mentions commit 6cbdd9726fb50. Before that commit, as I
mentioned in an earlier RFC, the LevelDB test is doing a lot worse in
mainline, which also matches what Tal tested in that cache_ext paper:
https://lore.kernel.org/linux-mm/20260502-mglru-fg-v1-0-913619b014d9@xxxxxxxxxxx/

It was like this:
GET/SCAN LevelDB [2]:
Classical LRU: throughput_avg: 993.64 Ops/s
MGLRU before: throughput_avg: 951.12 Ops/s
MGLRU after: throughput_avg: 1344.72 Ops/s (+41% faster)

After that commit, MGLRU beats CLRU for this test, but that's a side
effect. (And also I made the test run much longer to let both CLRU /
MGLRU to learn the hotness info more so in V2 RFC the ops/s is higher,
also reduced the RA window as RA window dramatically effects the
result).

6cbdd9726fb50 strongly favors the GET part, because of scan readahead:
non-readahead random folios (GET) will not use too many RA so they
mostly start with refs == 1 before landing in LRU (stay in percpu
folio batch, so also due to a time window and unstable) so they are
put in 2nd oldest gen. Sequential reads, however, use RA heavily and
start with refs == 0, landing in the oldest part. This creates a
strong bias toward reclaiming the scan part, which coincidentally
suits that benchmark. But it is just not right, assuming sequential
read is colder than non-sequential reads makes no sense. And it does
caused regressions for many other tests, e.g. some data base like
SQLite, and for example a recent report for other people:
https://lore.kernel.org/all/20260901180704.168106-1-ehab.ababneh@xxxxxxxxx/

With FG, however, it will actually learn the hotness. If you check
Tal's paper and code there is also warmup phase for any algorithm to
take effect, so I'm already glad that with NVMe, where the warmup
window is shorter, FG is already doing great, and even if the window
is long, there isn't significant regression, and more like a balance
issue. I believe this way will lead to better long-term performance.