Re: [RFC PATCH 0/4] memcg_ext: memcg policy through cgroup-attached struct_ops

From: Yafang Shao

Date: Thu Sep 24 2026 - 06:10:04 EST


On Wed, Sep 23, 2026 at 11:47 PM Shakeel Butt <shakeel.butt@xxxxxxxxx> wrote:
>
> Hi Yafang,
>
> On Wed, Sep 23, 2026 at 09:07:40PM +0800, Yafang Shao wrote:
> > On Tue, Sep 22, 2026 at 3:30 AM Shakeel Butt <shakeel.butt@xxxxxxxxx> wrote:
> > >
> [...]
> > >
> > > Known open questions
> > > ====================
> > >
> > > The semantics of memory.high for remote chargers or kernel threads is a
> > > grey area and this series does not aim to resolve that.
> > >
> > > Another open question is whether a bound on debt deferral is needed. At
> > > the moment, we think that rather than putting a limit on deferral for
> > > memory.high, it will be better to handle that through an async worker like
> > > memcg->high_work. We aim to introduce that later, along with the right CPU
> > > accounting for that async work.
> >
> > Hello Shakeel,
> >
> > On the open question of how deferred debt eventually gets paid: would
> > it make sense for the policy to also notify userspace (e.g. via
> > ringbuf) when it defers,
>
> I think the notification through bpf programs is already possible and a bpf
> program deciding to bypass memory.high can already do notification via ringbuf.
>
> > and have a userspace reclaimer do the reclaim
> > through memory.reclaim?
> >
> > I understand one of the concerns for the async worker is CPU
> > accounting. If the concern is that the kworker's CPU usage is not
> > charged to the target cgroup, the userspace reclaimer could instead be
> > spawned with clone3(CLONE_INTO_CGROUP) so it runs inside the target
> > cgroup, and both its CPU and memory usage get charged there.
> >
> > One caveat: intermediate cgroups with the no-internal-process
> > constraint cannot take processes, so this would only work for leaf
> > cgroups.
> >
> > What do you think?
>
> I think all of this is possible without additional code and with this series.
> With AI, should be very easy to prototype it. Please take a stab and I will look
> into it as well (time permitting).

An LLM helped me quickly implement a userspace async memcg reclaimer
based on your series, and it seems to work quite well.

>
> Thanks for taking a look and also please let me know what other ways you think
> memcg can be customized through BPF in a beneficial way.

Sure. On our production servers we have been running a set of BPF
programs to tailor kernel behavior for different workloads — all of
them global programs so far — and I believe they are all good
candidates for per-cgroup BPF policies now that cgroup-attached
struct_ops is available. They have been really helpful in our
Kubernetes production environment. I have sent some of them upstream,
such as:

- BPF-THP
https://lwn.net/Articles/1039689/
- BPF-auto-NUMA
https://lwn.net/Articles/1054030/

Perhaps we can revisit both of them and turn them into per-cgroup
policies — what do you think?

We are also running some custom BPF programs that have not been sent
upstream yet, such as:

- BPF-async-reclaimer
We don't care about the CPU accounting of the kworker, so we just
wake up a kworker to do the async reclaiming.
- BPF-fault-around

Both are really beneficial to our workloads, and both are global programs today.

We are planning a few more customizations to resolve painful
production issues, such as:

- The long-standing inode::lock contention caused by dentries [0].
We have not started implementing it yet, but we might introduce a
memcg->dentry_limit or a memcg->vfs_cache_pressure as BPF policies..
- cgroup-level readahead.

So, to answer your question directly: for memcg itself, the beneficial
customizations for us are the reclaim policy (the async reclaimer
above), the dentry/vfs cache pressure knobs, and fault-around; the
rest are per-cgroup MM policies that would need the
struct_ops-to-cgroup mechanism generalized beyond memcg — which is why
I hope these use cases can help make the design more generic.

[0] https://lore.kernel.org/linux-fsdevel/20240511200240.6354-2-torvalds@xxxxxxxxxxxxxxxxxxxx/

--
Regards
Yafang