[PATCH 2/3] tools/sched_ext: scx_pair: Drain task queue on cgroup exit

From: Wanwu Li

Date: Thu Sep 24 2026 - 10:44:11 EST


pair_cgroup_exit() releases a cgroup's queue slot by clearing the
busy marker and deleting the cgid hash entry, but leaves the task
queue (cgrp_q_arr) and its length counter (cgrp_q_len) alone. The
slot is handed to the next cgroup through pair_cgroup_init() without
any reset, so if the outgoing cgroup left entries behind -- tasks
which exited or migrated away after being enqueued, which scx_pair
never dequeues -- the inheriting cgroup starts with a non-zero
cgrp_q_len. pair_enqueue() only queues a cgroup on top_q on the
0 -> 1 transition of that counter, so the new cgroup never reaches
top_q and none of its tasks is ever dispatched: the cgroup starves
for the lifetime of the scheduler.

Fix it by draining the queue and taking the length counter to zero
with it, claiming each entry just as try_dispatch() does before
popping so that the drain cannot race an in-flight dispatcher into
a fatal pop failure. All pids left behind are stale so dropping
them is safe.

Fixes: f0262b102c7c ("tools/sched_ext: add scx_pair scheduler")
Signed-off-by: Wanwu Li <liwanwu@xxxxxxxxxx>
---
tools/sched_ext/scx_pair.bpf.c | 28 ++++++++++++++++++++++++++++
1 file changed, 28 insertions(+)

diff --git a/tools/sched_ext/scx_pair.bpf.c b/tools/sched_ext/scx_pair.bpf.c
index 0d61b7b812db..a1fe67fdaf5c 100644
--- a/tools/sched_ext/scx_pair.bpf.c
+++ b/tools/sched_ext/scx_pair.bpf.c
@@ -633,6 +633,34 @@ void BPF_STRUCT_OPS(pair_cgroup_exit, struct cgroup *cgrp)

q_idx = bpf_map_lookup_elem(&cgrp_q_idx_hash, &cgid);
if (q_idx) {
+ struct cgrp_q *cgq;
+ s32 pid;
+ u64 *cgq_len;
+
+ /*
+ * All tasks have left the cgroup by the time it exits, so
+ * the pids left behind are stale; drain the queue and take
+ * the length counter down with them -- the next cgroup
+ * inheriting this slot must see its first enqueue go
+ * 0 -> 1 or it never reaches top_q. Each pop claims a
+ * counter slot as try_dispatch() does, so that the drain
+ * can't steal an entry from an in-flight dispatcher and
+ * trip its scx_bpf_error().
+ */
+ cgq = bpf_map_lookup_elem(&cgrp_q_arr, q_idx);
+ cgq_len = MEMBER_VPTR(cgrp_q_len, [*q_idx]);
+ if (cgq && cgq_len)
+ bpf_repeat(BPF_MAX_LOOPS) {
+ u64 len = *(volatile u64 *)cgq_len;
+
+ if (!len)
+ break;
+ if (__sync_val_compare_and_swap(cgq_len, len, len - 1) != len)
+ continue;
+ if (bpf_map_pop_elem(cgq, &pid))
+ break;
+ }
+
u64 *busy = MEMBER_VPTR(cgrp_q_idx_busy, [*q_idx]);
if (busy)
*busy = 0;
--
2.25.1