Re: [RFC PATCH 0/4] memcg_ext: memcg policy through cgroup-attached struct_ops

From: Shakeel Butt

Date: Thu Oct 01 2026 - 19:02:08 EST


On Thu, Oct 01, 2026 at 08:42:55AM -1000, Tejun Heo wrote:
> Hello, Shakeel.
>
[...]
>
> > Here if you meant that default behavior of memory.high should work for most (if
> > not all) users then we are on same page. If some user want memory.high reclaim
> > to happen in a separate thread instead of return-to-userspace or synchronously,
> > this proposal provides mechanism through BPF to such users to achieve their
> > goals.
>
> I think there's a common reasonable solution here, which is deciding by who
> the charge is for. A task charging for itself gets return-to-userspace
> enforcement plus the checks in the explicit bulk operations above. A kthread
> or anything else charging on behalf of a cgroup shouldn't be throttled for
> the cgroup's overrun. The overrun is attributed to the cgroup, async reclaim
> is kicked, and the cgroup's own tasks absorb the throttling on their next
> return to userspace. The in-charge synchronous fallback goes away. Each case
> has one reasonable answer, so I don't see a policy choice to expose here.
>

Let me list the cases explicitly to see where we agree and where we disagree:

1. For the !in_task() charge path, today we trigger async reclaim and we will
continue to do the same in the future.

2. For a kthread (or remote charging), today we throttle it similarly to user
threads, but we want it to be handled similarly to the !in_task() case, i.e.
trigger async reclaim. Regarding your statement "the cgroup's own tasks
absorb the throttling on their next return to userspace", I assume you meant
that when some other user thread of that memcg goes through the charge path,
it will eventually do memory.high enforcement on return to userspace.

3. For a task, today we enforce memory.high on return to userspace, and if too
much charge is accumulated in a single kernel entry, we enforce the high
limit synchronously. You are suggesting that we remove the sync enforcement
and add a couple of throttling points at known bulk allocation sites.

Please correct me if I misunderstood.

We are in agreement on (1) and (2) completely. For (3), I am fine with removing
the sync enforcement, but for throttling points for bulk operation sites,
I think we should only add them when there is an actual use case for that
or someone complains about overrun from those sites.

Now, setting aside the default behavior of memory.high, I want to provide
additional flexibility to users for (3) specifically. One specific case is
letting users opt in to async reclaim instead of the other forms of memory.high
enforcement. Basically, users can specify that instead of having their
application threads throttled, they would prefer async reclaim to bring their
usage back below memory.high. Whether we provide this functionality through
BPF or through something else, I am open to options.

thanks,
Shakeel