Re: [RFC PATCH 0/4] memcg_ext: memcg policy through cgroup-attached struct_ops

From: Shakeel Butt

Date: Thu Sep 24 2026 - 16:45:55 EST


On Thu, Sep 24, 2026 at 06:01:27PM +0800, Yafang Shao wrote:
> On Wed, Sep 23, 2026 at 11:47 PM Shakeel Butt <shakeel.butt@xxxxxxxxx> wrote:
> >
> > Hi Yafang,
> >
> > On Wed, Sep 23, 2026 at 09:07:40PM +0800, Yafang Shao wrote:
> > > On Tue, Sep 22, 2026 at 3:30 AM Shakeel Butt <shakeel.butt@xxxxxxxxx> wrote:
> > > >
> > > Hello Shakeel,
> > >
> > > On the open question of how deferred debt eventually gets paid: would
> > > it make sense for the policy to also notify userspace (e.g. via
> > > ringbuf) when it defers,
> >
> > I think the notification through bpf programs is already possible and a bpf
> > program deciding to bypass memory.high can already do notification via ringbuf.
> >
> > > and have a userspace reclaimer do the reclaim
> > > through memory.reclaim?
> > >
> > > I understand one of the concerns for the async worker is CPU
> > > accounting. If the concern is that the kworker's CPU usage is not
> > > charged to the target cgroup, the userspace reclaimer could instead be
> > > spawned with clone3(CLONE_INTO_CGROUP) so it runs inside the target
> > > cgroup, and both its CPU and memory usage get charged there.
> > >
> > > One caveat: intermediate cgroups with the no-internal-process
> > > constraint cannot take processes, so this would only work for leaf
> > > cgroups.
> > >
> > > What do you think?
> >
> > I think all of this is possible without additional code and with this series.
> > With AI, should be very easy to prototype it. Please take a stab and I will look
> > into it as well (time permitting).
>
> An LLM helped me quickly implement a userspace async memcg reclaimer
> based on your series, and it seems to work quite well.

That's awesome. Please do take a look at the code and provide feedback and if
you don't mind, a tested-by tag would be awesome.

>
> >
> > Thanks for taking a look and also please let me know what other ways you think
> > memcg can be customized through BPF in a beneficial way.
>
> Sure. On our production servers we have been running a set of BPF
> programs to tailor kernel behavior for different workloads — all of
> them global programs so far — and I believe they are all good
> candidates for per-cgroup BPF policies now that cgroup-attached
> struct_ops is available. They have been really helpful in our
> Kubernetes production environment. I have sent some of them upstream,
> such as:
>
> - BPF-THP
> https://lwn.net/Articles/1039689/
> - BPF-auto-NUMA
> https://lwn.net/Articles/1054030/
>
> Perhaps we can revisit both of them and turn them into per-cgroup
> policies — what do you think?

Yes seems interesting and I remember other folks (I think Rik) were interested
in these ideas as well.

>
> We are also running some custom BPF programs that have not been sent
> upstream yet, such as:
>
> - BPF-async-reclaimer
> We don't care about the CPU accounting of the kworker, so we just
> wake up a kworker to do the async reclaiming.

I understand but I think for general solution we do need accounting for this and
I have rfc out for this.

> - BPF-fault-around
>
> Both are really beneficial to our workloads, and both are global programs today.
>
> We are planning a few more customizations to resolve painful
> production issues, such as:
>
> - The long-standing inode::lock contention caused by dentries [0].
> We have not started implementing it yet, but we might introduce a
> memcg->dentry_limit or a memcg->vfs_cache_pressure as BPF policies..
> - cgroup-level readahead.
>
> So, to answer your question directly: for memcg itself, the beneficial
> customizations for us are the reclaim policy (the async reclaimer
> above), the dentry/vfs cache pressure knobs, and fault-around; the
> rest are per-cgroup MM policies that would need the
> struct_ops-to-cgroup mechanism generalized beyond memcg — which is why
> I hope these use cases can help make the design more generic.
>

Thanks a lot for this information, I will think more on these.

> [0] https://lore.kernel.org/linux-fsdevel/20240511200240.6354-2-torvalds@xxxxxxxxxxxxxxxxxxxx/
>
> --
> Regards
> Yafang