Re: [PATCH v2 sched_ext/for-7.4] sched_ext: Keep proxy donors with slice left on the local DSQ

From: Andrea Righi

Date: Wed Oct 07 2026 - 03:16:11 EST


Hi Tejun,

On Tue, Oct 06, 2026 at 10:28:34AM -1000, Tejun Heo wrote:
...
> On Sat, Oct 03, 2026 at 12:15:58AM +0200, Andrea Righi wrote:
> > - if (p->scx.flags & SCX_TASK_IMMED) {
> > + if ((p->scx.flags & SCX_TASK_IMMED) && !p->is_blocked) {
> > p->scx.flags |= SCX_TASK_REENQ_PREEMPTED;
> > scx_do_enqueue_task(rq, p, SCX_ENQ_REENQ, -1);
>
> Sorry, I steered this the wrong way. Exempting donors from IMMED entirely
> overrides what the scheduler asked for on that task. An IMMED donor
> preempted by a higher class now sits on this CPU's local DSQ until the CPU
> gets back to it, where IMMED would have returned it to BPF to be placed
> where it's picked and resolved right away. The donor and the owner it's
> donating to end up waiting exactly where IMMED says they shouldn't.
>
> I think the condition we want is to keep an IMMED donor local only when
> it's about to be picked right away, which is the bookkeeping put from
> proxy_resched_idle(), and to treat it like any other IMMED task otherwise.
> That's what v1's proxy_put with PICK_PENDING did. Can we go back to that?

Ack, we can restore PICK_PENDING so an IMMED donor stays local only for the
bookkeeping put during proxy resolution. And a real preemption would return it
to BPF via ops.enqueue().

> The deferred scan then needs no blocked exemption: after the bookkeeping
> put, the donor is first and the rq is headed to idle, so the existing
> first && rq_is_open() test keeps it. The wakeup_preempt_scx() recheck
> isn't needed either, as a donor is never left where an unblocked IMMED
> task couldn't stay.

Agreed, we can remove the blocked-donor exemption from a deferred IMMED scan and
the extra wakeup recheck.

>
> > + if (p->scx.flags & SCX_TASK_IMMED)
> > + enq_flags |= SCX_ENQ_IMMED;
>
> Can you add a comment here? This reads as flag preservation, while the
> reason is that scx_caps_for_enq() maps IMMED to SCX_CAP_ENQ_IMMED, so a
> sub-sched holding only the base cap on the CPU can keep the donor local.

Ok.

>
> > - if (next && sched_class_above(&ext_sched_class, next->sched_class) &&
> > + if (!p->is_blocked &&
> > + next && sched_class_above(&ext_sched_class, next->sched_class) &&
> > scx_task_can_stay_on_cpu(rq, p)) {
>
> I suggested this but I don't think donors should be excluded here. For a
> scheduler without ENQ_LAST this never fires for a donor: dispatch_one()
> keeps it through KEEP_LAST and refills its slice. A scheduler with
> ENQ_LAST needs the signal on a donor as on any other last task: the CPU is
> going idle with the task still queued and BPF has to trigger the
> follow-up. Can you drop the !p->is_blocked?

Ok, makes sense, a blocked donor should receive SCX_ENQ_LAST when its scheduler
needs to arrange the follow-up scheduling event.

>
> The one put that changes is sched_proxy_block_task(), where
> proxy_reset_donor() puts the still-queued donor with the owner's
> execution context as @next. With a fair owner and a zeroed slice, that
> takes the LAST branch and the WARN fires for a scheduler without
> ENQ_LAST, on a legitimate path, so it needs handling along with the
> above. One idea, which may or may not work: proxy_needs_return() dequeues
> the donor before proxy_reset_donor() so that this put skips the QUEUED
> block. If sched_proxy_block_task() can do the same, dequeue_block_task()
> first, then the reset, then __block_task(), the put sees an unqueued task
> and the ops.enqueue() and ops.dequeue() pair the current order generates
> goes away too.

I think we can do this without modifying sched/core.c, sched_ext can set an
SCX_RQ_PROXY_BLOCKING flag around its calls to sched_proxy_block_task(). When
proxy_reset_donor() invokes put_prev_task_scx(), that flag tells sched_ext not
to reenqueue the donor and block_task() will dequeue it immediately afterward.
This should avoid the transient enqueue/dequeue pair and the false ENQ_LAST
warning.

This is separate from v1's SCX_RQ_PROXY_PICK_PENDING, which needs to be
re-introduced to distinguish the temporary proxy_resched_idle() put from a real
preemption of an IMMED donor. SCX_RQ_PROXY_BLOCKING, instead, identifies a donor
about to be blocked during a scheduler ownership change.

So we need to add two rq flags in this way, SCX_RQ_PROXY_PICK_PENDING and
SCX_RQ_PROXY_BLOCKING, but the whole logic stays in ext.c (with the flags
dfinition in sched.h). Does this approach makes sense to you?

Thanks,
-Andrea