Re: [RFC PATCH 0/4] memcg_ext: memcg policy through cgroup-attached struct_ops
From: Shakeel Butt
Date: Fri Oct 02 2026 - 18:20:05 EST
On Fri, Oct 02, 2026 at 06:25:08AM -1000, Tejun Heo wrote:
> Hello, Shakeel.
>
> On Thu, Oct 01, 2026 at 03:56:47PM -0700, Shakeel Butt wrote:
> > We are in agreement on (1) and (2) completely. For (3), I am fine with removing
> > the sync enforcement, but for throttling points for bulk operation sites,
> > I think we should only add them when there is an actual use case for that
> > or someone complains about overrun from those sites.
>
> On (3), if removing synchronous enforcement wouldn't regress anything,
> that's fine, but why was it added in the first place?
>
I added the sync enforcement to replace a Google internal feature which, on
memcg OOM, allows node controller couple of seconds to either increase the max
limit or let the memcg die. With sync enforcement, memory.high helped in simple
benchmarks. However later testing on some realistic Google workloads, I found
out that several thousand threads are very normal of typical Google workload and
memory.high sync enforcement is not effective on applications with large amount
of threads. In addition, there were workloads which on noticing blocked threads,
keep forking more threads. At the end implementing that feature using
memory.high didn't pan out.
> > Now, setting aside the default behavior of memory.high, I want to provide
> > additional flexibility to users for (3) specifically. One specific case is
> > letting users opt in to async reclaim instead of the other forms of memory.high
> > enforcement. Basically, users can specify that instead of having their
> > application threads throttled, they would prefer async reclaim to bring their
> > usage back below memory.high.
>
> As for flexibility, we already have a gradient of enforcement around
> memory.high. Is the need here to make the shape of that gradient
> configurable? Can you give specific examples where this is needed?
>
The concrete example I have is the kswapd like async reclaimers (plural) per
memcg. Kswapd is woken up on free pages falling below low watermark and then
when free pages fall below min watermark, allocators get throttled (enter direct
reclaim). I want to apply similar concept to memcg (but with right cpu
accounting and more concurrency).