Re: [RFC PATCH 0/4] memcg_ext: memcg policy through cgroup-attached struct_ops

From: Shakeel Butt

Date: Wed Sep 30 2026 - 09:50:52 EST


On Mon, Sep 28, 2026 at 10:40:20AM -1000, Tejun Heo wrote:
> Hello, Shakeel.
>
> On Mon, Sep 21, 2026 at 12:25:55PM -0700, Shakeel Butt wrote:
> > try_charge_memcg() calls __mem_cgroup_handle_over_high() before it returns,
> > which reclaims and can throttle the task. That happens wherever the charge
> > happens, so a task holding a kernel lock can be stuck there, and everything
> > waiting on that lock is stuck behind it.
>
> Why not just raise the lazy bound high enough that most charges never
> enforce inline, and maybe annotate the specific paths that can allocate a
> lot so that they do? Inline enforcement should be the exception, not the
> rule. Flipping that and then trying to reverse it with custom BPF policies
> doesn't make a lot of sense.
>

Please correct me if I misunderstood you. Mainly, you are saying that we
should have a sane default behavior for memory.high. At the moment, if the
current charging process accumulates charges totaling more than
MEMCG_CHARGE_BATCH pages and the target memcg is over its high limit,
memory.high is enforced synchronously. You are suggesting that we should
increase the threshold from MEMCG_CHARGE_BATCH to some arbitrarily large
number. In that case, synchronous enforcement of memory.high will be very
rare.

I am fine with changing the default behavior. Actually, I have been
contemplating whether I should propose a revert of commit c9afe31ec443e
("memcg: synchronously enforce memory.high for large overcharges") because
it has introduced more problems than it has solved, but that is a separate
topic. The initial commit already mentioned that MEMCG_CHARGE_BATCH was used
arbitrarily, so replacing it with something big might be acceptable. I want
to keep that decision separate.

I am not sure about annotating specific paths. I think it would impose a
greater maintenance burden as the kernel evolves, since the annotations
might become stale. Also, people might object to adding memcg-internal hooks
in non-memcg code paths. In any case, this can be explored separately.

Returning to the actual proposal, my plan was to start small with a narrow,
specific use case. However, my long-term plan is to provide a mechanism to
change the default behavior for custom use cases. For example, for
memory.high, I will provide a way for users to specify what behavior they
want, i.e., whether or not they want more synchronous throttling. Second, I
will introduce a mechanism to trigger async reclaim workers. I also plan to
extend this functionality to memory.max. I just wanted to convey that I will
keep pushing this proposal, with the use cases adjusted a bit.

> > One concrete scenario which can be resolved by this new feature is the
> > kernfs notify worker. It delivers notifications with the cgroup2
> > kernfs_rwsem held for read, and the charge for the delivery allocation goes
> > to the cgroup that set the watch, usually one already under pressure. So
> > the worker reclaims while holding the lock, a waiting writer blocks every
> > later reader, and anything touching cgroupfs stalls for seconds.
>
> Slowing down the kworker inline doesn't make sense. It's charging on behalf
> of the watcher through set_active_memcg(), which already tells us whose debt
> it is. Wouldn't it make more sense to defer the debt to that cgroup instead
> of slowing down the kernel thread?
>

Yes, that makes sense. When a non-task-context charge exceeds memory.high,
the kernel already schedules high_work for the charged memcg. I plan to do
the same for kthreads so they need not reclaim or throttle inline. This is
best-effort reclaim rather than exact debt accounting; I still need to
examine userspace tasks that charge another memcg through
set_active_memcg(). I am addressing the reclaim worker's CPU accounting
separately.

Thanks for taking a look and providing feedback.