Re: [RFC PATCH 3/4] memcg_ext: allow BPF to defer memory.high enforcement
From: Shakeel Butt
Date: Thu Sep 24 2026 - 17:25:05 EST
On Thu, Sep 24, 2026 at 01:14:59PM -0700, JP Kobryn wrote:
[...]
> > +static void bpf_memcg_ctx_init(struct bpf_memcg_ctx *ctx,
> > + struct mem_cgroup *memcg,
> > + struct mem_cgroup *over_limit, gfp_t gfp_mask)
> > +{
> > + ctx->memcg = memcg;
> > + ctx->memcg_over_limit = over_limit;
> > + ctx->task = current;
> > + ctx->cgroup_id = cgroup_id(memcg->css.cgroup);
> > + ctx->over_limit_cgroup_id = over_limit ?
> > + cgroup_id(over_limit->css.cgroup) : 0;
> > + ctx->nr_pages_over_high = current->memcg_nr_pages_over_high;
> > + ctx->gfp_flags = (__force u32)gfp_mask;
> > +}
> > +
> > +u32 bpf_memcg_high_policy(struct mem_cgroup *memcg,
> > + struct mem_cgroup *over_limit, gfp_t gfp_mask)
> > +{
> > + const struct bpf_prog_array_item *item;
> > + const struct bpf_memcg_ops *ops;
> > + struct bpf_memcg_ctx ctx;
> > + u32 acc = BPF_MEMCG_HIGH_NO_OPINION;
> > + struct cgroup *cgrp;
> > +
> > + if (!cgroup_bpf_enabled(CGROUP_MEMCG_OPS))
> > + return acc;
> > +
> > + /*
> > + * Only the default hierarchy has a cgroup_bpf, and the static key is
> > + * global, so one policy anywhere turns this on for v1 memcgs too. A
> > + * v1 memcg still cannot get here, because memory.high and swap.high
> > + * are both v2-only and so it never builds the debt that leads to this
> > + * call. A hook on a path v1 can reach needs its own cgroup_on_dfl()
> > + * test: a v1 cgroup has no effective array and an uninitialised
> > + * cgrp->bpf.refcnt.
> > + */
> > + cgrp = memcg->css.cgroup;
> > +
> > + /*
> > + * A program can allocate and re-enter the charge path. Skip the
> > + * nested call. This guards the callbacks only.
> > + */
> > + if (current->in_bpf_memcg)
> > + return acc;
> > + current->in_bpf_memcg = 1;
> > +
> > + rcu_read_lock_dont_migrate();
> > +
> > + /*
> > + * A memcg outlives its cgroup while it has charges, and
> > + * cgroup_bpf_release() frees the arrays when the cgroup goes.
> > + */
> > + if (!cgroup_bpf_tryget_live(cgrp))
> > + goto out;
> > +
> > + bpf_memcg_ctx_init(&ctx, memcg, over_limit, gfp_mask);
> > +
> > + bpf_cgroup_struct_ops_foreach(ops, item, cgrp, CGROUP_MEMCG_OPS) {
> > + if (ops->high_policy)
> > + acc |= ops->high_policy(&ctx) &
> > + BPF_MEMCG_HIGH_VALID_MASK;
> > + }
>
> If I'm reading correctly, the gfp_mask at this point doesn't account for
> task restrictions, so the BPF program may see __GFP_FS, __GFP_IO, etc
> which may later be cleared when setting up the scan_control instance.
I am passing the same gfp_mask which has been passed to try_charge_memcg(), so
it should be same as what reclaim_high (reclaim internal) sees.