[RFC PATCH] mm: vmscan: avoid over-reclaim in classic LRU direct reclaim path

From: Bo Zhang

Date: Thu Oct 08 2026 - 08:02:05 EST


From: Bo Zhang <zhangbo56@xxxxxxxxxx>

On production Android devices, direct reclaim is a frequent source of
unpredictable, large latency spikes. A direct reclaim request typically
asks for only nr_to_reclaim = SWAP_CLUSTER_MAX (32) pages, and in the
common case completes in well under a millisecond. But the classic LRU
path has no reliable "goal met" check, so under memory pressure a single
reclaim can scan tens of thousands of pages and free hundreds to
thousands - far beyond the 32 requested - before returning.

Pooling the 20 slowest direct-reclaim events, before this patch the
worst case scanned 130,079 pages and reclaimed 1,635, and the average
scanned 37,955 pages and reclaimed 881.

The root cause is over-reclaim: the classic LRU path keeps scanning well
past its target in two ways:

1. shrink_lruvec() does not stop as soon as nr_to_reclaim is met; it
keeps scanning the remaining LRU batches, over-reclaiming past the
target.

2. shrink_node_memcgs() only breaks early on a partial walk. When the
initial partial walk misses nr_to_reclaim, do_try_to_free_pages()
retries with a full walk, which has no "goal met" check and shrinks
every memcg in the tree even after the target is satisfied.

Add an early exit to shrink_lruvec() and shrink_node_memcgs(), and the
matching check to should_continue_reclaim(), all keyed on

sc->nr_reclaimed >= max(sc->nr_to_reclaim, compact_gap(sc->order))

The compact_gap() term keeps correctness for high-order reclaim, where
nr_to_reclaim alone may be too small to free a block of the requested
order; MGLRU's should_abort_scan() uses the same threshold. Using it in
all three places keeps the retry decision consistent with where reclaim
actually stops.

After this patch, over the same 20 slowest events the worst case scans
45,593 pages and reclaims 57, and the average scans 6,399 pages and
reclaims 38 - reclaim now stays close to the requested target instead
of running far past it. Removing this over-reclaim greatly reduces the
direct-reclaim tail latency, which significantly cuts the jank caused by
direct reclaim on interactive workloads.

Signed-off-by: Bo Zhang <zhangbo56@xxxxxxxxxx>
---
mm/vmscan.c | 35 +++++++++++++++++++++++++++++++++--
1 file changed, 33 insertions(+), 2 deletions(-)

diff --git a/mm/vmscan.c b/mm/vmscan.c
index 91295070ca33..b1efbd2a9bb6 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -6236,6 +6236,18 @@ static void shrink_lruvec(struct lruvec *lruvec, struct scan_control *sc)

cond_resched_tasks_rcu_qs();

+ /*
+ * For high-order allocations, nr_to_reclaim may be smaller
+ * than what is needed to make a block of the requested order
+ * available for compaction. When memory is plentiful, keep
+ * reclaiming until at least compact_gap() pages have been
+ * freed so the target order can actually be satisfied.
+ */
+ if (!cgroup_reclaim(sc) &&
+ sc->nr_reclaimed + nr_reclaimed >=
+ max(nr_to_reclaim, compact_gap(sc->order)))
+ break;
+
if (nr_reclaimed < nr_to_reclaim || proportional_reclaim)
continue;

@@ -6345,6 +6357,16 @@ static inline bool should_continue_reclaim(struct pglist_data *pgdat,
if (!nr_reclaimed)
return false;

+ /*
+ * Stop once the overall reclaim goal is met, using the same
+ * max(nr_to_reclaim, compact_gap()) threshold as the early exits
+ * in shrink_lruvec() and shrink_node_memcgs(), so the retry
+ * decision here stays consistent with where reclaim actually
+ * stops.
+ */
+ if (sc->nr_reclaimed >= max(sc->nr_to_reclaim, compact_gap(sc->order)))
+ return false;
+
/* If compaction would go ahead or the allocation would succeed, stop */
for_each_managed_zone_pgdat(zone, pgdat, z, sc->reclaim_idx) {
unsigned long watermark = min_wmark_pages(zone);
@@ -6442,8 +6464,17 @@ static void shrink_node_memcgs(pg_data_t *pgdat, struct scan_control *sc)
sc->nr_scanned - scanned,
sc->nr_reclaimed - reclaimed);

- /* If partial walks are allowed, bail once goal is reached */
- if (partial && sc->nr_reclaimed >= sc->nr_to_reclaim) {
+ /*
+ * Break loop once the overall reclaim goal is met. For
+ * high-order allocations nr_to_reclaim may be too small to
+ * free a block of the requested order, so when memory is
+ * plentiful keep going until at least compact_gap() pages
+ * have been reclaimed. This applies to both partial and full
+ * walks: a full walk is only a means to reach the goal, not a
+ * requirement to visit every memcg.
+ */
+ if (sc->nr_reclaimed >= max(sc->nr_to_reclaim,
+ compact_gap(sc->order))) {
mem_cgroup_iter_break(target_memcg, memcg);
break;
}
--
2.34.1