Re: [RFC PATCH 0/4] memcg_ext: memcg policy through cgroup-attached struct_ops
From: Yafang Shao
Date: Sun Sep 27 2026 - 23:24:01 EST
On Fri, Sep 25, 2026 at 4:42 AM Shakeel Butt <shakeel.butt@xxxxxxxxx> wrote:
>
> On Thu, Sep 24, 2026 at 06:01:27PM +0800, Yafang Shao wrote:
> > On Wed, Sep 23, 2026 at 11:47 PM Shakeel Butt <shakeel.butt@xxxxxxxxx> wrote:
> > >
> > > Hi Yafang,
> > >
> > > On Wed, Sep 23, 2026 at 09:07:40PM +0800, Yafang Shao wrote:
> > > > On Tue, Sep 22, 2026 at 3:30 AM Shakeel Butt <shakeel.butt@xxxxxxxxx> wrote:
> > > > >
> > > > Hello Shakeel,
> > > >
> > > > On the open question of how deferred debt eventually gets paid: would
> > > > it make sense for the policy to also notify userspace (e.g. via
> > > > ringbuf) when it defers,
> > >
> > > I think the notification through bpf programs is already possible and a bpf
> > > program deciding to bypass memory.high can already do notification via ringbuf.
> > >
> > > > and have a userspace reclaimer do the reclaim
> > > > through memory.reclaim?
> > > >
> > > > I understand one of the concerns for the async worker is CPU
> > > > accounting. If the concern is that the kworker's CPU usage is not
> > > > charged to the target cgroup, the userspace reclaimer could instead be
> > > > spawned with clone3(CLONE_INTO_CGROUP) so it runs inside the target
> > > > cgroup, and both its CPU and memory usage get charged there.
> > > >
> > > > One caveat: intermediate cgroups with the no-internal-process
> > > > constraint cannot take processes, so this would only work for leaf
> > > > cgroups.
> > > >
> > > > What do you think?
> > >
> > > I think all of this is possible without additional code and with this series.
> > > With AI, should be very easy to prototype it. Please take a stab and I will look
> > > into it as well (time permitting).
> >
> > An LLM helped me quickly implement a userspace async memcg reclaimer
> > based on your series, and it seems to work quite well.
>
> That's awesome. Please do take a look at the code and provide feedback and if
> you don't mind, a tested-by tag would be awesome.
For this series, feel free to add:
Tested-by: Yafang Shao <laoar.shao@xxxxxxxxx>
One nit:
Since we might extend the cgroup-ops to other subsystems, or even to
the cgroup itself, could we define a more generic CGROUP_OPS instead
of CGROUP_MEMCG_OPS, and also rename `struct bpf_memcg_ops` to `struct
bpf_cgroup_ops`? This way, we attach to the `struct cgroup` itself
rather than to a specific subsystem.
--
Regards
Yafang