Re: [PATCH] mm: vmscan: don't count per-node proactive reclaim as memory pressure
From: Barry Song
Date: Tue Oct 06 2026 - 06:18:54 EST
On Mon, Oct 5, 2026 at 6:21 PM Vernon Yang <vernon2gm@xxxxxxxxx> wrote:
>
> On Mon, Oct 05, 2026 at 02:30:36PM +0800, Barry Song wrote:
> > On Mon, Oct 5, 2026 at 2:23 PM Vernon Yang <vernon2gm@xxxxxxxxx> wrote:
> > >
> > > From: Vernon Yang <yanglincheng@xxxxxxxxxx>
> > >
> > > When the cgroup has no memory pressure at all, writing to
> > > /sys/devices/system/nodeX/reclaim triggers proactive reclaim on
> > > NUMA node, causing increase in the writer cgroup's memory PSI.
> > >
> > > Due to this reclaim is performed in the context of the write(),
> > > accounted as memory pressure on the writer, like
> > > commit e22c6ed90aa9 ("mm: memcontrol: don't count limit-setting reclaim
> > > as memory pressure"). This is unexpected, the phenomenon resembling
> > > senpai will appear again.
> > >
> > > The Documentation/ABI/stable/sysfs-devices-node documentation also
> > > notes that "This interface is equivalent to the memcg variant."
> > >
> > > This patch unifies the semantics of the memcg and node interfaces:
> > > per-node proactive reclaim is no longer counted as memory pressure,
> > > and the per-node proactive reclaim interface no longer produces
> > > phantom pressure.
> > >
> > > I ran demo[1] that performs per-node proactive reclaim 10000 times
> > > in qemu, writer cgroup memory.pressure as follows:
> > >
> > > without patch:
> > >
> > > some avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
> > > full avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
> > >
> > > with patch:
> > >
> > > some avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
> > > full avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
> > >
> > > [1] https://github.com/vernon2gh/app_and_module/tree/main/reclaim_node_psi
> > >
> > > Fixes: b980077899ea ("mm: introduce per-node proactive reclaim interface")
> > > Cc: stable@xxxxxxxxxxxxxxx
> > > Signed-off-by: Vernon Yang <yanglincheng@xxxxxxxxxx>
> > > ---
> >
> > This seems to be a valid concern. Personally, I don't like
> > having the code depend on whether `__node_reclaim()` and
> > `node_reclaim()` are called from proactive reclaim or page
> > allocation.
> >
> > Can't we check whether `sc->proactive` is true? Am I missing
> > something?
>
> It is also fine to directly check `sc->proactive` in __node_reclaim().
>
> I chose the current coding because a previous similar fix commit
> e22c6ed90aa9 was written this way, and it is also very clear.
>
> Of course, it depends on everyone's preference. If everyone prefers to
> directly check `sc->proactive`, please let me know clearly. Thanks!
I personally think the following is much more explicit than the
current implicit code based on path dependency. no?
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 6a931b576b7d..048ba11852fe 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -7997,7 +7997,8 @@ static unsigned long __node_reclaim(struct
pglist_data *pgdat,
sc->gfp_mask);
cond_resched();
- psi_memstall_enter(&pflags);
+ if (!sc->proactive)
+ psi_memstall_enter(&pflags);
delayacct_freepages_start();
fs_reclaim_acquire(sc->gfp_mask);
/*
@@ -8014,7 +8015,8 @@ static unsigned long __node_reclaim(struct
pglist_data *pgdat,
memalloc_noreclaim_restore(noreclaim_flag);
fs_reclaim_release(sc->gfp_mask);
delayacct_freepages_end();
- psi_memstall_leave(&pflags);
+ if (!sc->proactive)
+ psi_memstall_leave(&pflags);
trace_mm_vmscan_node_reclaim_end(sc->nr_reclaimed, NULL);
>
> --
> Cheers,
> Vernon