Re: [PATCH RFC] sched_ext: warn when cpu.max is set but the BPF scheduler doesn't implement bandwidth control
From: Tao Cui
Date: Tue Aug 18 2026 - 09:59:59 EST
Hi, tejun
在 2026/8/18 21:53, Tao Cui 写道:
> From: Tao Cui <cuitao@xxxxxxxxxx>
>
> The kernel stores cpu.max bandwidth parameters in the task_group and
> passes them to the BPF scheduler via ops.cgroup_set_bandwidth() and
> scx_cgroup_init_args, but does not enforce the quota itself. If the
> loaded BPF scheduler doesn't implement the callback, cpu.max is
> silently ignored -- the cgroup gets unlimited CPU regardless of the
> configured quota.
>
> Of the example schedulers, only scx_qmap implements the callback --
> and only to bpf_printk() the parameters, so no in-tree scheduler
> actually enforces the quota. Measured with scx_simple: a
> cgroup with cpu.max = "50000 100000" (50% of one CPU) and one
> busy task used 9946ms of CPU in 10 seconds with nr_throttled
> remaining 0.
>
> Print a one-time warning when a finite quota is configured on a
> cgroup while the active scheduler lacks the callback, so users and
> container orchestrators know the quota is not enforced.
>
Some background on how I found this: I was testing how sched_ext
interacts with cgroup CPU controls in a VM, and the cpu.max case stood
out:
cgroup with cpu.max = "50000 100000" (50% of one CPU),
one busy task, scx_simple loaded
sched_ext : 9946ms of CPU in 10s, nr_throttled = 0
CFS : ~5000ms in 10s, throttling as expected
The same happens with scx_flatcg and scx_central -- neither they nor
scx_simple implement ops.cgroup_set_bandwidth(), so the quota is
silently ignored. grep shows scx_qmap is the only in-tree user of the
callback -- and its implementation just bpf_printk()s the parameters, so
even there the quota is not enforced. Outside the
tree, lavd implements its own bandwidth accounting, but as far as I can
tell rusty and bpfland don't, which suggests users on those schedulers
are running containers with cpu.max that does nothing.
That's why I drafted the warning patch -- but I'm not sure a warning is
the right approach. Some alternatives I can think of:
1. pr_warn_once() as in this patch (minimal, but the quota is still
not enforced)
2. refuse to enable sched_ext (or the cgroup support) when a finite
quota exists and the callback is missing
3. kernel-side fallback enforcement, e.g. throttle in
scx_next_task_picked()/dispatch path based on tg->scx.bw_quota_us
Is the missing enforcement considered the BPF scheduler's responsibility
by design (and just under-documented), or would a kernel fallback be
welcome? If it's the former, maybe sched-ext.rst should mention that
cpu.max requires ops.cgroup_set_bandwidth() from the loaded scheduler.
Happy to work on whichever direction you prefer.
Thanks,
Tao
> Signed-off-by: Tao Cui <cuitao@xxxxxxxxxx>
> ---
> kernel/sched/ext/ext.c | 6 ++++++
> 1 file changed, 6 insertions(+)
>
> diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
> index 10af28a9f2c0..de786d0b928a 100644
> --- a/kernel/sched/ext/ext.c
> +++ b/kernel/sched/ext/ext.c
> @@ -4953,6 +4953,12 @@ void scx_group_set_bandwidth(struct task_group *tg,
> tg->scx.bw_burst_us != burst_us))
> SCX_CALL_OP(sch, cgroup_set_bandwidth, NULL,
> tg_cgrp(tg), period_us, quota_us, burst_us);
> + else if (scx_cgroup_enabled && sch &&
> + !SCX_HAS_OP(sch, cgroup_set_bandwidth) &&
> + quota_us != RUNTIME_INF)
> + pr_warn_once("sched_ext: BPF scheduler \"%s\" does not implement "
> + "ops.cgroup_set_bandwidth(); cpu.max will not be enforced\n",
> + sch->ops.name);
>
> tg->scx.bw_period_us = period_us;
> tg->scx.bw_quota_us = quota_us;