Re: [PATCH v5 sched_ext/for-7.3 28/33] sched_ext: Route ops.update_idle() to sub-schedulers and re-notify owed scheds

From: Andrea Righi

Date: Tue Jul 14 2026 - 02:19:33 EST


Hi Tejun,

On Thu, Jul 09, 2026 at 12:50:36PM -1000, Tejun Heo wrote:
> __scx_update_idle() notified only the root scheduler. A sub-scheduler that
> holds a cid needs that cid's idle state to place and kick on it.
>
> Deliver ops.update_idle() to every scheduler that holds SCX_CAP_BASE on the
> transitioning cid. The root holds every cap, so a real transition always
> reaches it.
>
> Real transitions are not enough on their own. A cid that is already idle
> when a sub-sched gains baseline access produces no transition, so the new
> holder would never learn it is idle. The ecaps sync arms a re-notify on the
> gain, and the next idle pick delivers ops.update_idle() to just that sched,
> leaving holders that already track the cpu untouched. A matching loss of
> baseline access drops any pending re-notify.
>
> Bypass suppresses ops.update_idle() too, so a cpu that goes idle during a
> bypass window and stays idle yields no transition to re-deliver on
> un-bypass. Arm the same re-notify for every sched leaving bypass. The acute
> case is a child granted cids during its own ops.sub_attach(). The grant
> lands while the child is bypassed and the notify walk skips it, so on
> un-bypass it holds cids it never saw go idle. The root is owed the same and
> is armed through a separate per-rq flag, which keeps this working when
> sub-schedulers are compiled out.
>
> v2: Gate the idle catch-up in pick_task_idle() to avoid a double ops.update_idle(). (sashiko AI)
>
> Signed-off-by: Tejun Heo <tj@xxxxxxxxxx>
> ---

...

> +/*
> + * Notify schedulers of an idle transition on @cpu's cid, delivering to every
> + * sched that holds %SCX_CAP_BASE on the cid (the root holds every cap). A real
> + * transition (@do_notify) reaches all holders. A forced one (@root_renotify for
> + * the root, a sub-sched's idle_renotify marker for a sub) reaches only the owed
> + * scheds.
> + */
> +static void scx_idle_notify(struct rq *rq, bool idle, bool do_notify, bool root_renotify)
> +{
> + s32 cpu = cpu_of(rq);
> + s32 cid = scx_cpu_arg(cpu);
> + struct scx_sched *pos;
> +
> + lockdep_assert_rq_held(rq);
> +
> + pos = scx_next_descendant_pre(NULL, scx_root);
> + while (pos) {
> + bool forced = false;
> +
> + if (unlikely(scx_missing_caps(pos, cpu, SCX_CAP_BASE))) {
> + pos = scx_skip_subtree_pre(pos, scx_root);
> + continue;
> + }
> +
> + if (pos == scx_root) {
> + forced = root_renotify;
> + }
> +#ifdef CONFIG_EXT_SUB_SCHED
> + else if (per_cpu_ptr(pos->pcpu, cpu)->idle_renotify) {
> + per_cpu_ptr(pos->pcpu, cpu)->idle_renotify = false;
> + forced = true;
> + }
> +#endif
> + if ((do_notify || forced) && SCX_HAS_OP(pos, update_idle) &&
> + !scx_bypassing(pos, cpu))
> + SCX_CALL_OP(pos, update_idle, rq, cid, idle);
> + pos = scx_next_descendant_pre(pos, scx_root);
> + }
> +}

So, this makes every real idle/busy transition walk the whole scheduler
hierarchy and potentially invoke ops.update_idle() for every cap-holding
scheduler while the rq lock is held and IRQs are disabled?

This becomes O(number of cap-holding schedulers) BPF callbacks per idle
transition.

I haven't benchmarked this, so I'm not sure if it's a valid performance concern.
We don't have to fix this now, it can be a future improvement. In that case, if
we prove that we have a real bottleneck here, would it make sense to maintain a
per-rq list of schedulers that both hold effective SCX_CAP_BASE and implement
ops.update_idle()?

Also, for the root-only case, would it be worth keeping the old direct
ops.update_idle() path behind a static key which is enabled while any
sub-scheduler is attached?

Thanks,
-Andrea