[PATCH 0/3] tools/sched_ext: Fix three example scheduler bugs

From: Wanwu Li

Date: Thu Sep 24 2026 - 10:50:57 EST


Patch 1 fixes scx_userland, where the BPF component never increments
nr_queued although both its declaration comment and the NOTE in
userland_update_idle() say it does. As a result the "wake the
user-space scheduler when a CPU goes idle" decision silently degrades
to nr_scheduled alone and tasks sitting in the enqueued map are
invisible to it.

Patches 2 and 3 fix two related queue-lifetime bugs in scx_pair, which
has no .dequeue and no .cgroup_move callback and therefore leaves stale
pids in the per-cgroup FIFOs:

- pair_cgroup_exit() releases a queue slot without draining it. The
next cgroup inheriting the slot starts with a stale non-zero
cgrp_q_len, so its first enqueue never sees the 0 -> 1 transition
and its cgid is never queued on top_q: the cgroup starves for the
lifetime of the scheduler.

- try_dispatch() trusts a pid popped from a cgroup's FIFO. If the
task migrated to another cgroup between enqueue and dispatch, it is
dispatched under the old cgroup's context, violating the core
"paired CPUs only run tasks from the same cgroup" invariant.

Wanwu Li (3):
tools/sched_ext: Increment nr_queued when a task is queued to user
space
tools/sched_ext: scx_pair: Drain task queue on cgroup exit
tools/sched_ext: scx_pair: Verify task cgroup before dispatch

tools/sched_ext/scx_pair.bpf.c | 47 ++++++++++++++++++++++++++++++++++++++
tools/sched_ext/scx_userland.bpf.c | 1 +
2 files changed, 48 insertions(+)

--
2.25.1