Re: [PATCH] mm: vmscan: don't count per-node proactive reclaim as memory pressure
From: Vernon Yang
Date: Tue Oct 06 2026 - 08:32:45 EST
On Tue, Oct 06, 2026 at 06:17:02PM +0800, Barry Song wrote:
> On Mon, Oct 5, 2026 at 6:21 PM Vernon Yang <vernon2gm@xxxxxxxxx> wrote:
> >
> > On Mon, Oct 05, 2026 at 02:30:36PM +0800, Barry Song wrote:
> > > On Mon, Oct 5, 2026 at 2:23 PM Vernon Yang <vernon2gm@xxxxxxxxx> wrote:
> > > >
> > > > From: Vernon Yang <yanglincheng@xxxxxxxxxx>
> > > >
> > > > When the cgroup has no memory pressure at all, writing to
> > > > /sys/devices/system/nodeX/reclaim triggers proactive reclaim on
> > > > NUMA node, causing increase in the writer cgroup's memory PSI.
> > > >
> > > > Due to this reclaim is performed in the context of the write(),
> > > > accounted as memory pressure on the writer, like
> > > > commit e22c6ed90aa9 ("mm: memcontrol: don't count limit-setting reclaim
> > > > as memory pressure"). This is unexpected, the phenomenon resembling
> > > > senpai will appear again.
> > > >
> > > > The Documentation/ABI/stable/sysfs-devices-node documentation also
> > > > notes that "This interface is equivalent to the memcg variant."
> > > >
> > > > This patch unifies the semantics of the memcg and node interfaces:
> > > > per-node proactive reclaim is no longer counted as memory pressure,
> > > > and the per-node proactive reclaim interface no longer produces
> > > > phantom pressure.
> > > >
> > > > I ran demo[1] that performs per-node proactive reclaim 10000 times
> > > > in qemu, writer cgroup memory.pressure as follows:
> > > >
> > > > without patch:
> > > >
> > > > some avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
> > > > full avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
> > > >
> > > > with patch:
> > > >
> > > > some avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
> > > > full avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
> > > >
> > > > [1] https://github.com/vernon2gh/app_and_module/tree/main/reclaim_node_psi
> > > >
> > > > Fixes: b980077899ea ("mm: introduce per-node proactive reclaim interface")
> > > > Cc: stable@xxxxxxxxxxxxxxx
> > > > Signed-off-by: Vernon Yang <yanglincheng@xxxxxxxxxx>
> > > > ---
> > >
> > > This seems to be a valid concern. Personally, I don't like
> > > having the code depend on whether `__node_reclaim()` and
> > > `node_reclaim()` are called from proactive reclaim or page
> > > allocation.
> > >
> > > Can't we check whether `sc->proactive` is true? Am I missing
> > > something?
> >
> > It is also fine to directly check `sc->proactive` in __node_reclaim().
> >
> > I chose the current coding because a previous similar fix commit
> > e22c6ed90aa9 was written this way, and it is also very clear.
> >
> > Of course, it depends on everyone's preference. If everyone prefers to
> > directly check `sc->proactive`, please let me know clearly. Thanks!
>
> I personally think the following is much more explicit than the
> current implicit code based on path dependency. no?
Yes, both implementations are fine with me.
I usually wait 1~2 weeks to give all maintainer/reviewer to review.
As long as no one raises objections during that period, all suggestions
will be accepted in the next version, don't worry. Thanks for your
suggestions!
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index 6a931b576b7d..048ba11852fe 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -7997,7 +7997,8 @@ static unsigned long __node_reclaim(struct
> pglist_data *pgdat,
> sc->gfp_mask);
>
> cond_resched();
> - psi_memstall_enter(&pflags);
> + if (!sc->proactive)
> + psi_memstall_enter(&pflags);
> delayacct_freepages_start();
> fs_reclaim_acquire(sc->gfp_mask);
> /*
> @@ -8014,7 +8015,8 @@ static unsigned long __node_reclaim(struct
> pglist_data *pgdat,
> memalloc_noreclaim_restore(noreclaim_flag);
> fs_reclaim_release(sc->gfp_mask);
> delayacct_freepages_end();
> - psi_memstall_leave(&pflags);
> + if (!sc->proactive)
> + psi_memstall_leave(&pflags);
>
> trace_mm_vmscan_node_reclaim_end(sc->nr_reclaimed, NULL);
>