Re: [PATCH v4 6/6] mm/mglru: fix potential generation folio number leak
From: Kairui Song
Date: Tue Sep 01 2026 - 07:54:52 EST
On Tue, Sep 1, 2026 at 6:56 PM Barry Song <baohua@xxxxxxxxxx> wrote:
>
> On Mon, Aug 31, 2026 at 2:43 AM Kairui Song via B4 Relay
> <devnull+kasong.tencent.com@xxxxxxxxxx> wrote:
> >
> > From: Kairui Song <kasong@xxxxxxxxxxx>
> >
> > Each generation of MGLRU accounts anon and file folio numbers
> > separately, and the page table walker updates each generation's
> > counters in batch once the walk is done. The walker promotes a
> > folio's generation with a cmpxchg on folio->flags, and
> > update_batch_size() then reads the live flags again to pick the
> > anon/file column to charge. The walk holds neither the lruvec lock nor
> > the folio lock, so the type can flip between the cmpxchg and that
> > read: the lazyfree path clears PG_swapbacked, and reclaim sets it back
> > on a dirty lazyfree folio. The batched delta pair is then recorded in
> > the wrong type column. Nothing reconciles it afterwards, permanently
> > skewing lrugen->nr_pages and the reclaim budgets derived from it.
> >
> > Fix it by capturing the type from the flags snapshot the cmpxchg
> > linearized against: folio_update_gen() returns the type of the state
> > it transitioned from, and update_batch_size() accounts with that.
> >
> > A folio's type only changes while it is off the LRU list, inside a
> > del/add pair under the lruvec lock, with the gen bits cleared in
> > between. The generation and PG_swapbacked sit in the same
> > folio->flags word, so the cmpxchg snapshot captures them together.
> > Let G be the generation that snapshot captured (old_gen) and G' the
> > one it wrote (new_gen); the CAS can land in only three places:
> >
> > - before the del: the folio is anon at G; the batch records anon
> > G -> G', and the del later removes the folio from the anon
> > counters;
> > - between del and add: gen == -1, so folio_update_gen() returns -1
> > without touching the flags and no batch is recorded; the del/add
> > pair accounts for the move alone;
> > - after the add: the folio is file at the fresh generation the add
> > charged; the batch records file, that gen -> G', matching that
> > charge.
> >
> > Unlike the drift of lazy promotions, which sort_folio() repairs under
> > the lruvec lock, the phantom deltas from before this fix land in a
> > column the folio never occupies again, so nothing ever repairs them.
> >
> > Fixes: 018ee47f1489 ("mm: multi-gen LRU: exploit locality in rmap")
> > Signed-off-by: Kairui Song <kasong@xxxxxxxxxxx>
> > ---
> > include/linux/mm_inline.h | 7 ++++++-
> > mm/vmscan.c | 13 +++++++------
> > 2 files changed, 13 insertions(+), 7 deletions(-)
> >
Thanks for the review!
> > diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h
> > index 047295ae6e8a..7e487c23aff7 100644
> > --- a/include/linux/mm_inline.h
> > +++ b/include/linux/mm_inline.h
> > @@ -10,6 +10,11 @@
> > #include <linux/userfaultfd_k.h>
> > #include <linux/leafops.h>
> >
> > +static inline int folio_flags_is_file_lru(const unsigned long *flags)
> > +{
> > + return !test_bit(PG_swapbacked, flags);
> > +}
> > +
> > /**
> > * folio_is_file_lru - Should the folio be on a file LRU or anon LRU?
> > * @folio: The folio to test.
> > @@ -27,7 +32,7 @@
> > */
> > static inline int folio_is_file_lru(const struct folio *folio)
> > {
> > - return !folio_test_swapbacked(folio);
> > + return folio_flags_is_file_lru(const_folio_flags(folio, 0));
> > }
>
> I guess we don't need to change this function? It seems more natural to
> me to keep using `!folio_test_swapbacked(folio)` here.
>
> In `folio_update_gen()`, you already use
> `*is_file = folio_flags_is_file_lru(&old_flags)` to get the type from
> `old_flags`. I guess that's all we need?
Yeah, that's all we need, I was thinking that using
folio_flags_is_file_lru here can avoid any further change to
PG_swapbacked causing inconsistency of folio_is_file_lru and
folio_flags_is_file_lru, seems a bit easier to maintain.
>
> >
> > static __always_inline void __update_lru_size(struct lruvec *lruvec,
> > diff --git a/mm/vmscan.c b/mm/vmscan.c
> > index 724b6e034e69..87e667c410ed 100644
> > --- a/mm/vmscan.c
> > +++ b/mm/vmscan.c
> > @@ -3269,7 +3269,8 @@ static bool positive_ctrl_err(struct ctrl_pos *sp, struct ctrl_pos *pv)
> > ******************************************************************************/
> >
> > /* promote pages accessed through page tables */
> > -static int folio_update_gen(struct folio *folio, int new_gen, const vma_flags_t *vma_flags)
> > +static int folio_update_gen(struct folio *folio, int new_gen, int *is_file,
>
> Rather than is_file, I feel was_file would be more accurate here,
> since it doesn't represent what the folio is now.
I'm fine either way, is_file seems more natual to me. Or maybe simply
rename it as "type"? Just the below:
> > @@ -3532,9 +3533,9 @@ static void walk_update_folio(struct lru_gen_mm_walk *walk, struct vm_area_struc
> > folio_mark_dirty(folio);
> >
> > if (walk) {
> > - old_gen = folio_update_gen(folio, new_gen, &vma->flags);
> > + old_gen = folio_update_gen(folio, new_gen, &file, &vma->flags);
> > if (old_gen >= 0 && old_gen != new_gen)
> > - update_batch_size(walk, folio, old_gen, new_gen);
> > + update_batch_size(walk, folio, old_gen, new_gen, file);
>
> This is one more place where we rely on `LRU_GEN_ANON = 0` and
> `LRU_GEN_FILE = 1`.
>
> We do this kind of thing quite often in MGLRU, such as `type = !type`.
> It relies on the numeric values of the type constants, which is sort of
> a semantic abuse, but it's hard to find a shorter way to express it.
We also ahve WORKINGSET_ANON, WORKINGSET_FILE, ANON_AND_FILE. I can
rename the newly added parameter to "type" here if that looks better.