RE: [RFC PATCH] mm/mglru: dynamically protect readahead fault folios under refault pressure
From: Ababneh, Ehab
Date: Thu Oct 01 2026 - 13:22:46 EST
Hi Kairui,
> -----Original Message-----
> From: Kairui Song <ryncsn@xxxxxxxxx>
> Sent: Saturday, September 26, 2026 12:06 AM
> To: Ababneh, Ehab <ehab.ababneh@xxxxxxxxx>
> Cc: Barry Song <baohua@xxxxxxxxxx>; akpm@xxxxxxxxxxxxxxxxxxxx; Axel
> Rasmussen <axelrasmussen@xxxxxxxxxx>; Lance Yang
> <lance.yang@xxxxxxxxx>; LKML <linux-kernel@xxxxxxxxxxxxxxx>; linux-mm
> <linux-mm@xxxxxxxxx>; Qi Zheng <qi.zheng@xxxxxxxxx>; Shakeel Butt
> <shakeel.butt@xxxxxxxxx>; Wei Xu <weixugc@xxxxxxxxxx>; Yuanchu Xie
> <yuanchu@xxxxxxxxxx>; Yu Zhao <yuzhao@xxxxxxxxxx>
> Subject: Re: [RFC PATCH] mm/mglru: dynamically protect readahead fault
> folios under refault pressure
>
> Ababneh, Ehab <ehab.ababneh@xxxxxxxxx> 于 2026年9月25日周五
> 01:54写道:
> >
> > Hi Barry,
> >
> > > -----Original Message-----
> > > From: Barry Song <baohua@xxxxxxxxxx>
> > > Sent: Wednesday, September 23, 2026 9:08 PM
> > > To: Ababneh, Ehab <ehab.ababneh@xxxxxxxxx>
> > > Cc: ryncsn@xxxxxxxxx; akpm@xxxxxxxxxxxxxxxxxxxx;
> > > axelrasmussen@xxxxxxxxxx; kasong@xxxxxxxxxxx;
> lance.yang@xxxxxxxxx;
> > > linux-kernel@xxxxxxxxxxxxxxx; linux-mm@xxxxxxxxx;
> > > qi.zheng@xxxxxxxxx; shakeel.butt@xxxxxxxxx; weixugc@xxxxxxxxxx;
> > > yuanchu@xxxxxxxxxx; yuzhao@xxxxxxxxxx
> > > Subject: Re: [RFC PATCH] mm/mglru: dynamically protect readahead
> > > fault folios under refault pressure
> > >
> > > On Thu, Sep 24, 2026 at 11:43 AM Ababneh, Ehab
> > > <ehab.ababneh@xxxxxxxxx> wrote:
> > > >
> > > > Hi Kairui, Barry,
> > > >
> > > > Apologies for the delay in getting back to you with these results
> > > > — my test environment got corrupted and I had to spend some time
> > > > recovering it before I could re-run the benchmark.
> > >
> > > No worries, Ehab. And thanks very much for your testing!
> > >
> > > [...]
> > >
> > > > > > > I agree that reverting the commit that caused the regression
> > > > > > > is not the optimal path. I expect there are many workloads
> > > > > > > and scenarios that benefit from the behavior introduced by
> > > > > > > that commit, so reverting it could unnecessarily regress those
> workloads.
> > > > > > >
> > > > > > > I will run the Cassandra benchmark with Kairui's MGLRU-FG
> > > > > > > patches to see whether they address the issue I am seeing. I
> > > > > > > will send the results when they are ready.
> > > > > > >
> > > >
> > > > Thanks for the pointer — I gave your MGLRU-FG fix a try against
> > > > the same Cassandra read benchmark I used for my patch (4 nodes,
> > > > 720s, 100
> > > readers).
> > > >
> > > > Results with MGLRU-FG:
> > > > Op rate: 55.0k - 56.5k op/s
> >
> > >
> > > I noticed this is even higher than your dynamic readahead fix:
> > > "throughput ~51.9k-52.7k op/s"
> > >
> > > Is this a run-to-run variation, or is the improvement stable across runs?
> > >
> >
> > Yes, I can't explain the slight gain in throughput but loss in latency
> > compared to my fix. I'll probably study this in more detail when I can
> > dedicate some time to it.
>
> Hi Ehab,
>
> That's very interesting, I think it means we are getting a bit better at identifying
> the working set. But maybe a slightly high cost of aging? And P99 is usually
> very sensitive to noises, one long tailing aging event could make it look a lot
> worse.
>
> > Typically, I get a unique result immediately after a reboot, but
> > subsequent runs are very consistent. The variation is usually around
> > 0.1 ms gain/loss in latency and approximately 100-200 requests per
> > second in throughput. I discard the results from the first run.
>
> Any way I can reproduce this? E.g. what is the test bench, database size,
> benchmark, machine spec, etc?
>
> > > > Latency 99th percentile: 6.8 - 6.9 ms
> > > >
> > > > For reference, here's where the other variants landed on the same setup:
> > > >
> > > > Regression (6cbdd9726fb5, "mm/mglru: use folio_mark_accessed to
> > > > replace folio_set_active"):
> > > > p99 ~9.2-9.5 ms, throughput ~41.8k-43.6k op/s
> > > >
> > > > Revert of 6cbdd9726fb5:
> > > > p99 ~5.5-5.6 ms, throughput ~51.9k-53.1k op/s
> > > >
> > > > My dynamic readahead-credit fix:
> > > > p99 ~5.8 ms, throughput ~51.9k-52.7k op/s
> > > >
> > > > So MGLRU-FG recovers most of the latency regression, but the
> > > > revert and my fix still recover more of it: p99 with MGLRU-FG is
> > > > noticeably higher than with the revert and my fix (6.8-6.9 ms vs
> > > > ~5.5-5.8 ms), though it's still a big improvement over the
> > > > regression's 9.2-9.5 ms
>
> Did you test the V2 or V1 of MGLRU-FG? V2 includes Barry's aging
> optimization; I'm not sure if this latency is caused by aging or maybe some
> other change.
>
> And is there any swap device you used, or is the test using the same base
> commit? MM also changes significantly between different baselines.
My setup runs Cassandra 4.1.5 on Java 11. The host has two Intel Xeon
Platinum 8468 sockets, 96 physical CPU cores (192 logical CPUs), two
NUMA nodes, and 440 GiB of system-visible RAM. Cassandra instances
use CPU affinity, with memory interleaved across NUMA nodes.
I run four Cassandra instances, each backed by a dedicated 3.5 TB
NVMe drive. Each drive has a separate mount, and holds its instance’s
data files, commit log, and saved caches.
The stress workload uses a pre-populated table. Its schema, partitioning,
replication, compaction, and compression settings, along with the stress
profile and concurrency, are recorded because they can affect performance.
I used more than one kernel version across my test runs. The most recent
was Linux 7.3.0-rc5. Within each set of setups being compared, I kept the
kernel version consistent.
For the MGLRU-FG tests, I used the ryncsn/b4/mglru-fg-v1.8 branch. The
MGLRU-FG implementation is commit 6bcf4ae70b5e; the branch tip
is a619ab0105d9.