[PATCH 2/3] sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ
From: Kuba Piecuch
Date: Tue Sep 29 2026 - 12:18:07 EST
When the BPF scheduler moves a task to the local DSQ of a CPU other than
the one whose rq the task is on, e.g. with SCX_DSQ_LOCAL_ON dispatch,
scx_bpf_dsq_move_to_local() or scx_bpf_dsq_move*(), the task is migrated
by move_remote_task_to_local_dsq(). p->scx.sticky_cpu is set across the
migration so that dequeueing the task from the source rq isn't treated as
the task leaving the BPF scheduler's custody.
However, enqueue_task_scx() only clears p->scx.sticky_cpu after
scx_do_enqueue_task() has inserted the task into the destination local
DSQ. task_leave_custody(), called from rq_owned_post_enq() on insertion,
still sees the migration in progress and skips the custody exit. The task
ends up on a terminal DSQ with SCX_TASK_IN_CUSTODY set and without
ops.dequeue() having been called.
The custody exit then happens only when the task is picked for execution,
in set_next_task_scx(), which reports it to the BPF scheduler as
ops.dequeue(SCX_DEQ_CORE_SCHED_EXEC) even though no core-sched pick took
place. If a scheduling property change hits the task while it's waiting
on the local DSQ, the BPF scheduler instead gets
ops.dequeue(SCX_DEQ_SCHED_CHANGE) for a task that has already left its
custody. Both break the ops.dequeue() semantics, under which a dispatch to
a terminal DSQ ends custody with an ops.dequeue() call without special
flags.
Clear p->scx.sticky_cpu before calling scx_do_enqueue_task(). The routing
decision in scx_do_enqueue_task() uses the local copy, and the departure
side is unaffected as p->scx.sticky_cpu is still set across
deactivate_task(). ops.dequeue() is now invoked on the destination rq when
the task is inserted into the local DSQ, as it already is for same-rq
dispatches.
Fixes: ebf1ccff79c4 ("sched_ext: Fix ops.dequeue() semantics")
Cc: stable@xxxxxxxxxxxxxxx # v7.1+
Assisted-by: Claude:claude-opus-5.5
Signed-off-by: Kuba Piecuch <jpiecuch@xxxxxxxxxx>
---
kernel/sched/ext/ext.c | 16 +++++++++++-----
1 file changed, 11 insertions(+), 5 deletions(-)
diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index ad391a8cbd05..763de7797056 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -1501,8 +1501,9 @@ static inline bool task_scx_migrating(struct task_struct *p)
/*
* We only need to check sticky_cpu: it is set to the destination
* CPU in move_remote_task_to_local_dsq() before deactivate_task()
- * and cleared when the task is enqueued on the destination, so it
- * is only non-negative during an internal SCX migration.
+ * and cleared in enqueue_task_scx() on the destination before @p is
+ * inserted into the local DSQ, so it is only non-negative while @p
+ * is in transit between the two rqs.
*/
return p->scx.sticky_cpu >= 0;
}
@@ -2189,10 +2190,15 @@ static void enqueue_task_scx(struct rq *rq, struct task_struct *p, int core_enq_
if (rq->scx.nr_running == 1)
dl_server_start(&rq->ext_server);
- scx_do_enqueue_task(rq, p, enq_flags, sticky_cpu);
+ /*
+ * An SCX-internal migration ends once @p arrives on the destination
+ * rq. Clear sticky_cpu before enqueueing so that @p leaves the BPF
+ * scheduler's custody when inserted into the destination local DSQ.
+ * The local copy in @sticky_cpu is used for routing.
+ */
+ p->scx.sticky_cpu = -1;
- if (sticky_cpu >= 0)
- p->scx.sticky_cpu = -1;
+ scx_do_enqueue_task(rq, p, enq_flags, sticky_cpu);
out:
rq->scx.flags &= ~SCX_RQ_IN_WAKEUP;
--
2.56.0.rc1.315.gc6ed9934b7-goog