Re: [PATCH] mm: mglru: clear the reference counter for rejected folios

From: Baolin Wang

Date: Tue Sep 08 2026 - 06:18:37 EST




On 9/8/26 4:23 PM, Baoquan He wrote:
On 09/08/26 at 03:53pm, Baolin Wang wrote:


On 9/8/26 2:59 PM, Baoquan He wrote:
On 09/08/26 at 12:01pm, Baolin Wang wrote:


On 9/8/26 11:03 AM, Baolin Wang wrote:


On 9/8/26 10:34 AM, Barry Song wrote:
On Tue, Sep 8, 2026 at 10:30 AM Baoquan He <baoquan.he@xxxxxxxxx> wrote:

Hi Baolin,

On 09/07/26 at 11:25am, Baolin Wang wrote:
......snip...
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 40d3f1b48a74..42c0a09938ab 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c

Well, this seems to be based on Andrew's mm-new branch. I usually track
mm-unstable branch. Maybe the subject should be marked as below?
[PATCH mm-new] mm: mglru: clear the reference counter for rejected

ACK.


@@ -5021,10 +5021,11 @@ static int evict_folios(unsigned
long nr_to_scan, struct lruvec *lruvec,
               }

               /* don't add rejected folios to the oldest generation */
-             if (lru_gen_folio_seq(lruvec, folio, false) ==
min_seq[type]) {
-                     folio_set_lru_refs(folio, 0);
+             if (lru_gen_folio_seq(lruvec, folio, false) ==
min_seq[type])
                       folio_set_active(folio);
-             }
+
+             /* See the comments on LRU_REFS_FLAGS */
+             folio_set_lru_refs(folio, 0);

This looks like a great catch, while the code change could bring issue.

Because move_folios_to_lru() relies on folios' flags to decide their new
generation. You just cleared it before move_folios_to_lru(). This is no
problem for rejected folios that are determined to be put into the
oldest generation. But for those rejected folios that are determined to
be promoted, this could be wrong. E.g currently gen window is 4, and a
folio is referenced, lru_gen_folio_seq() decides its new gen as 1, which
is the 2nd oldest generation. While folio_set_lru_refs(folio, 0) clear
referenced bit, this causes it being put into the oldest generation in
move_folios_to_lru(), this is not expected.

Yes. As I discussed with Barry earlier, lru_gen_folio_seq() also needs
to be reconsidered regarding whether it should rely on PG_referenced
[1].

For commit 6cbdd9726fb5, we didn't discuss the impact on rejected folios
either. Before commit 6cbdd9726fb5, if rejected folios did not have
PG_active set by shrink_folio_list(), evict_folios() would set PG_active
on these rejected folios.

[1] https://lore.kernel.org/linux-mm/20260901220430.79810-1-
baohua@xxxxxxxxxx/

The original code looks quite weird. It even prioritizes folios
that won't be promoted by `PG_active`. Do we need to change all
the cases just to call `PG_active`?

I would agree if we can. While I have one concern. A rejected folio that
should have gone into the oldest generation is now being promoted to the
2nd newest generation. In the original code, referenced folio is only
being promoted to the next gen. I even think this is not a bug, but Yu

That's not quite true. Before commit 6cbdd9726fb5, a rejected referenced
folio was also put back to the 2nd youngest gen.

OK, I didn't follow your earlier discussion, I need take some time to
fully understand that commit and the patch thread from Ehab.


Zhao intentionally did it: the coldest folio is moved to 2nd newest gen,
referenced folio (hot folio) is moved to new gen but carries the referenced
bit. Both of them seems to be treated somewhat equally.

To me, I would rather move both of them to the next gen, while keep their
refs untouched.

IMHO, I strongly recommend not doing this, and that's exactly the motivation
behind my patch. Because this is also being treated as a promotion, it
should behave like folio_inc_gen() or folio_update_gen() and clear the refs
after promotion. I think this is a fundamental principle of promotion.
Otherwise, ref-based promotion is already completely broken.

OK, it makes sense to me to make principle of promotion strictly applied
no efficiency degradation involved.


Next, I also plan to clean up refs in lru_gen_set_refs() as discussed with
Barry.

Looks forward to seeing that. By the way, your discussion with Baryr is
private or in public list, do you have pointer if public? Thanks.

I raised this issue before[1], and recently there has been another discussion[2] about it (we had some private discussions, but I'll post a new patch for discussion).

[1] https://lore.kernel.org/all/eb395442-0aad-428a-a5ac-9072d2d89060@xxxxxxxxxxxxxxxxx/
[2] https://lore.kernel.org/linux-mm/20260901220430.79810-1-baohua@xxxxxxxxxx/

Yes. Regarding this concern, I plan to change back to the original
behavior:

    /* See the comments on LRU_REFS_FLAGS */
    folio_set_lru_refs(folio, 0);

    /* don't add rejected folios to the oldest generation */
    if (lru_gen_folio_seq(lruvec, folio, false) == min_seq[type])
        folio_set_active(folio);

What do you think?

Just FYI, after above changes, the performance improvement on zram is no
longer obvious either. I think this also answers Kairui's earlier question
about why I saw a performance improvement (which seems related to commit
6cbdd9726fb5). Also, there is no obvious performance regression either.

Hi Baolin,

Not sure if it's convenient to do a little more testing in your side.
E.g rejected folios are moved to next gen, but not clearing their flags.

I'm not sure what you mean by "next gen" here. If you mean the 2nd oldest
gen, I actually tested that too, that is, clearing refs after
move_folios_to_lru(), and there wasn't any noticeable performance impact
either (but the code was a bit hacky, so I didn't go with this approach).

Yeah, I meant the 2nd oldest gen, while what I am curious about is moving
them into 2nd oldest gen but not clearing refs. Clearly it's conflicting
with your plan. Anyway, it's just a brain store idea, please forget it.
Thanks for the sharing and detailed explanation.

Thanks for reviewing.