Re: [PATCH] memcg: trim the per-cpu charge stock instead of draining it
From: Shakeel Butt
Date: Tue Aug 18 2026 - 11:12:49 EST
On Tue, Aug 18, 2026 at 12:03:07PM +0200, Michal Hocko wrote:
> On Mon 17-08-26 16:46:51, Shakeel Butt wrote:
> > Joy reported that an application generating a request/response traffic
> > pattern spends 44.6% to 57.0% of CPU in the memcg charge/uncharge path
> > for a range of message sizes, against 0.27% to 0.71% outside that range.
> > Running from the root memcg, where socket memory accounting is skipped,
> > recovers the performance.
> >
> > Tracing the charge path showed that the application generates a pattern
> > where the write syscall charges one page and the read syscall uncharges
> > two pages on the same CPU. This hits a corner case in the memcg percpu
> > stock code that thrashes the stock continuously.
> >
> > In the memcg percpu stock code, MEMCG_CHARGE_BATCH (64) is both the high
> > watermark and the emptying target, i.e. on a request to charge one page
> > the kernel charges MEMCG_CHARGE_BATCH pages and caches
> > (MEMCG_CHARGE_BATCH - 1) of them in the percpu stock. The following
> > uncharge of 2 pages takes the cached count to (MEMCG_CHARGE_BATCH + 1),
> > and refill_stock() then empties the cache completely. With such a
> > pattern the percpu stock becomes completely ineffective.
> >
> > Instead of a single boundary point for charges, use the technique the
> > page allocator uses for its own percpu caches, which keeps the watermark
> > and the emptying target apart: nr_pcp_free() frees between batch and
> > high - batch pages, leaving at least pcp->batch on the list. Add a high
> > watermark MEMCG_STOCK_HIGH and, once the cached count goes over it,
> > return only the pages above MEMCG_STOCK_LOW. The watermarks are
> > MEMCG_CHARGE_BATCH apart, so a page_counter update still covers a full
> > batch. Peak cached pages per memcg grows from 64 to 96, the same
> > high-versus-batch tradeoff the page allocator makes.
>
> The idea is sound. I would just not increase the overall stock size in
> the same patch. Fine tuning can be done independently and ideally with
> some numbers.
> Would it make sense to start with MEMCG_STOCK_HIGH := MEMCG_CHARGE_BATCH
> and MEMCG_CHARGE_BATCH := MEMCG_CHARGE_BATCH / 2. That would preserve
> the maximum stock size while preventing all or nothing behavior which is
> indeed suboptimal and pushing charging path to a slower path way too
> aggressively.
>
> WDYT?
Yes, this makes sense. Let me run the experiment with that workload to make sure
the newer number works and resend the patch.
Thanks for the review.