Re: [PATCH sched_ext/for-7.2-fixes] sched_ext: mark waker CPU busy when selected in WAKE_SYNC case
From: Andrea Righi
Date: Sun Jul 26 2026 - 02:41:24 EST
Hi Kuba,
On Wed, Jul 22, 2026 at 02:33:07PM +0000, Kuba Piecuch wrote:
> SCX's built-in idle CPU tracking is imperfect and can be out-of-sync
> with the actual idle state of CPUs, especially immediately after
> enabling SCX due to scx_idle_enable() marking all online CPUs idle.
>
> scx_select_cpu_dfl() skips marking the selected CPU as busy if
> the selected CPU is the waker CPU in the SCX_WAKE_SYNC case.
> If the waker CPU was marked idle by SCX, it will still be marked idle
> after CPU selection and potentially even after switching to the wakee.
>
> In the allowed_cpus selftest, this can manifest as the test failing with
> the following message in dmesg:
>
> allowed_cpus.bpf.c:21: CPU 0 should be marked as busy
>
> This patch explicitly marks the waker CPU as busy. With this patch,
> the test failure no longer reproduces. There are still some pretty
> unlikely races that could make the test fail (e.g. pick_task_idle()
> marking the selected CPU idle between selection and validation), but
> these can't be fixed easily.
>
> Signed-off-by: Kuba Piecuch <jpiecuch@xxxxxxxxxx>
Thanks for fixing this. I'm wondering if we should also rework the allowed_cpus
selftest, since it's inevitably racy and can trigger false positives.
In particular, we should avoid checking whether a remotely selected CPU is still
absent from the idle mask after scx_bpf_select_cpu_and() returns. The helper
only claims the CPU at selection time, and an idle-to-idle re-pick can
legitimately set it back to idle again before the BPF program performs the
check.
But the logic should work for the local CPU in ops.select_cpu(): if
scx_bpf_cpu_curr() reports a non-idle scheduling context, the local CPU must not
be present in the idle mask. This should be race-free because the local CPU
cannot run pick_task_idle() while the callback is executing.
And we can also run a short bootstrap phase, forcing a task to run on each
online CPU, to make sure the idle cpumasks are properly initialised.
This would make the selftest validate stable properties rather than a transient
state. I'll send a patch in a bit.
Thanks,
-Andrea
> ---
> kernel/sched/ext/idle.c | 4 +++-
> 1 file changed, 3 insertions(+), 1 deletion(-)
>
> diff --git a/kernel/sched/ext/idle.c b/kernel/sched/ext/idle.c
> index 8e8c6201b7df..6f93cc32b650 100644
> --- a/kernel/sched/ext/idle.c
> +++ b/kernel/sched/ext/idle.c
> @@ -554,8 +554,10 @@ s32 scx_select_cpu_dfl(struct task_struct *p, s32 prev_cpu, u64 wake_flags,
> cpu_rq(cpu)->scx.local_dsq.nr == 0 &&
> (!(flags & SCX_PICK_IDLE_IN_NODE) || (waker_node == node)) &&
> !cpumask_empty(idle_cpumask(waker_node)->cpu)) {
> - if (cpumask_test_cpu(cpu, allowed))
> + if (cpumask_test_cpu(cpu, allowed)) {
> + scx_idle_test_and_clear_cpu(cpu);
> goto out_unlock;
> + }
> }
> }
>
> --
> 2.55.0.229.g6434b31f56-goog
>