Re: [PATCH V3 3/3] mm: page_alloc: drain pcp lists before oom kill
From: Yosry Ahmed
Date: Fri Sep 04 2026 - 12:47:09 EST
On Fri, Sep 4, 2026 at 9:27 AM Michal Hocko <mhocko@xxxxxxxx> wrote:
>
> On Fri 04-09-26 09:09:35, Yosry Ahmed wrote:
> > On Fri, Sep 4, 2026 at 4:21 AM Michal Hocko <mhocko@xxxxxxxx> wrote:
> > >
> > > On Thu 03-09-26 07:03:33, Yosry Ahmed wrote:
> > > > On Thu, Sep 3, 2026 at 12:27 AM Michal Hocko <mhocko@xxxxxxxx> wrote:
> > > > >
> > > > > On Wed 02-09-26 16:49:48, Yosry Ahmed wrote:
> > > > > > On Sun, Nov 05, 2023 at 06:20:50PM +0530, Charan Teja Kalla wrote:
> > > > > > > pcp lists are drained from __alloc_pages_direct_reclaim(), only if some
> > > > > > > progress is made in the attempt.
> > > > > > >
> > > > > > > struct page *__alloc_pages_direct_reclaim() {
> > > > > > > .....
> > > > > > > *did_some_progress = __perform_reclaim(gfp_mask, order, ac);
> > > > > > > if (unlikely(!(*did_some_progress)))
> > > > > > > goto out;
> > > > > > > retry:
> > > > > > > page = get_page_from_freelist();
> > > > > > > if (!page && !drained) {
> > > > > > > drain_all_pages(NULL);
> > > > > > > drained = true;
> > > > > > > goto retry;
> > > > > > > }
> > > > > > > out:
> > > > > > > }
> > > > > > >
> > > > > > > After the above, allocation attempt can fallback to
> > > > > > > should_reclaim_retry() to decide reclaim retries. If it too return
> > > > > > > false, allocation request will simply fallback to oom kill path without
> > > > > > > even attempting the draining of the pcp pages that might help the
> > > > > > > allocation attempt to succeed.
> > > > > > >
> > > > > > > VM system running with ~50MB of memory shown the below stats during OOM
> > > > > > > kill:
> > > > > > > Normal free:760kB boost:0kB min:768kB low:960kB high:1152kB
> > > > > > > reserved_highatomic:0KB managed:49152kB free_pcp:460kB
> > > > > > >
> > > > > > > Though in such system state OOM kill is imminent, but the current kill
> > > > > > > could have been delayed if the pcp is drained as pcp + free is even
> > > > > > > above the high watermark.
> > > > > > >
> > > > > > > Fix this missing drain of pcp list in should_reclaim_retry() along with
> > > > > > > unreserving the high atomic page blocks, like it is done in
> > > > > > > __alloc_pages_direct_reclaim().
> > > > > > >
> > > > > > > Signed-off-by: Charan Teja Kalla <quic_charante@xxxxxxxxxxx>
> > > > > >
> > > > > > [Sorry for thread necromancy]
> > > > > >
> > > > > > Hi Charan,
> > > > > >
> > > > > > Are you planning to respin this patch?
> > > > > >
> > > > > > I know that Michal was questioning the need for it. While doing some
> > > > > > stress testing I came across a couple of OOM kills that had significant
> > > > > > amount of memory in pcplists. Something that would have been prevented
> > > > > > by this patch.
> > > > >
> > > > > Could you share some numbers to see the scale of the problem?
> > > >
> > > > Sure, here's a sample from an OOM log (ignore mapped/free_mapped, it's
> > > > from the ALLOC_UNMAPPED series):
> > > >
> > > > [ 40.336188] Mem-Info:
> > > > [ 40.336195] active_anon:96 inactive_anon:1540181 isolated_anon:0
> > > > active_file:51 inactive_file:0 isolated_file:0
> > > > unevictable:382720 dirty:43 writeback:0
> > > > slab_reclaimable:20339 slab_unreclaimable:26597
> > > > mapped:382858 shmem:153 pagetables:8291
> > > > sec_pagetables:0 bounce:0
> > > > kernel_misc_reclaimable:0
> > > > free:17264 free_pcp:16890 free_cma:0
> > > > [ 40.336199] Node 0 active_anon:384kB inactive_anon:6160724kB
> > > > active_file:204kB inactive_file:0kB unevictable:1530880kB
> > > > isolated(anon):0kB isolated(file):0kB mapped:1531432kB dirty:172kB
> > > > writeback:0kB shmem:612kB shmem_thp:0kB shmem_pmdmapped:0kB
> > > > anon_thp:10240kB kernel_stack:2576kB pagetables:33164kB
> > > > sec_pagetables:0kB all_unreclaimable? yes Balloon:0kB gpu_active:0kB
> > > > gpu_reclaim:0kB
> > > > [ 40.336202] DMA32 free:34808kB boost:0kB min:11028kB low:13784kB
> > > > high:16540kB reserved_highatomic:0KB free_highatomic:0KB
> > > > free_mapped:11032KB active_anon:0kB inactive_anon:1487720kB
> > > > active_file:52kB inactive_file:20kB unevictable:393252kB
> > > > writepending:4kB zspages:0kB present:2096760kB managed:1991156kB
> > > > mlocked:393252kB bounce:0kB free_pcp:42404kB local_pcp:2460kB
> > > > free_cma:0kB
> > > > [ 40.336205] lowmem_reserve[]: 0 5976 5976
> > > > [ 40.336210] Normal free:34248kB boost:0kB min:34024kB low:42528kB
> > > > high:51032kB reserved_highatomic:0KB free_highatomic:0KB
> > > > free_mapped:34028KB active_anon:384kB inactive_anon:4672764kB
> > > > active_file:180kB inactive_file:0kB unevictable:1137628kB
> > > > writepending:168kB zspages:0kB present:6291456kB managed:6120144kB
> > > > mlocked:1137628kB bounce:0kB free_pcp:25156kB local_pcp:764kB
> > > > free_cma:0kB
> > > > [ 40.336213] lowmem_reserve[]: 0 0 0
> > > >
> > > > As far as I can tell there's about ~65M of free memory stranded on
> > > > pcplists (in an 8G VM), which would have kept the amount of free
> > > > memory above the watermarks and prevented that specific OOM kill.
> > > > Although, as I mentioned before, this is a synthetic stress test.
> > > > Perhaps in practice it doesn't matter all that much in practice, but
> > > > it seems like the logical thing to do.
> > >
> > > yes, pcp lists are quite (unusually) high. Is it possible they simply
> > > got repopulated after the first direct reclaim run? Or is there
> > > something else(odd) going on?
> >
> > I think in this case there wasn't a lot of reclaimable memory (mostly
> > anon with no swap), so at some point direct reclaim didn't make any
> > progress. __alloc_pages_direct_reclaim() only drains the pcplists if
> > reclaim makes progress, I am not sure if that's intentional. An
> > alternative would be perhaps removing this condition, something like
> > this:
> >
> > diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> > index 8d79f76cdd0e1..98e9079240ad5 100644
> > --- a/mm/page_alloc.c
> > +++ b/mm/page_alloc.c
> > @@ -4592,11 +4592,12 @@ __alloc_pages_direct_reclaim(gfp_t gfp_mask,
> > unsigned int order,
> > psi_memstall_enter(&pflags);
> > *did_some_progress = __perform_reclaim(gfp_mask, order, ac);
> > if (unlikely(!(*did_some_progress)))
> > - goto out;
> > + goto drain;
> >
> > retry:
> > page = get_page_from_freelist(gfp_mask, order, alloc_flags, ac);
> >
> > +drain:
> > /*
> > * If an allocation failed after direct reclaim, it could be because
> > * pages are pinned on the per-cpu lists or in high alloc reserves.
> > @@ -4608,7 +4609,6 @@ __alloc_pages_direct_reclaim(gfp_t gfp_mask,
> > unsigned int order,
> > drained = true;
> > goto retry;
> > }
> > -out:
> > psi_memstall_leave(&pflags);
>
> Ideally if we can make the function call less hairy. Maybe we want to
> make draining part of the reclaim as the last resort when normal reclaim
> fails.
Do you mean do the draining in __perform_reclaim(), or deeper into the
reclaim stack?
The thing is that __alloc_pages_direct_reclaim() currently drains when
__perform_reclaim() fails to make any progress and we still cannot
allocate. The change above makes it drain if it cannot allocate after
__perform_reclaim(), regardless of progress. So if you want to move it
into __perform_reclaim(), we'll have it in both places.
Or maybe I just don't understand what you meant :)