Re: [RFC PATCH 0/4] memcg_ext: memcg policy through cgroup-attached struct_ops

From: Tejun Heo

Date: Thu Oct 01 2026 - 14:43:09 EST


Hello, Shakeel.

On Wed, Sep 30, 2026 at 06:00:41PM -0700, Shakeel Butt wrote:
> mlock(), madvise(POPULATE), fadvise(WILLNEED) are the obvious ones where large
> amount of memory can be allocated before returning to userspace.

These are explicit bulk memory operations running in the calling task's own
context, so they can be special-cased. Each is a loop with points between
chunks where no locks are held, and the same over-high handler the
return-to-userspace path runs can run there. That keeps the throttling on
the task which asked for the memory and can't cause priority inversions. The
set of such paths is small and obvious, so I don't think it'd be onerous to
maintain.

> > This doesn't seem like a policy problem.
>
> What is "This" in the above statement?

The whole thing, how memory.high should be enforced when the overrun happens
inside the kernel or from a kthread.

> Here if you meant that default behavior of memory.high should work for most (if
> not all) users then we are on same page. If some user want memory.high reclaim
> to happen in a separate thread instead of return-to-userspace or synchronously,
> this proposal provides mechanism through BPF to such users to achieve their
> goals.

I think there's a common reasonable solution here, which is deciding by who
the charge is for. A task charging for itself gets return-to-userspace
enforcement plus the checks in the explicit bulk operations above. A kthread
or anything else charging on behalf of a cgroup shouldn't be throttled for
the cgroup's overrun. The overrun is attributed to the cgroup, async reclaim
is kicked, and the cgroup's own tasks absorb the throttling on their next
return to userspace. The in-charge synchronous fallback goes away. Each case
has one reasonable answer, so I don't see a policy choice to expose here.

Thanks.

--
tejun