Re: [PATCH 07/10] sched_ext: Add proxy destination query kfuncs

From: Andrea Righi

Date: Sat Jul 11 2026 - 05:07:45 EST


On Fri, Jul 10, 2026 at 02:54:39PM -0700, John Stultz wrote:
> On Fri, Jul 10, 2026 at 1:40 AM Andrea Righi <arighi@xxxxxxxxxx> wrote:
> >
> > BPF schedulers admitting blocked proxy donors may want to know the CPU
> > or cid where the mutex owner will execute.
> >
>
> As I mentioned before, this probably needs some more context as to why
> this is useful/important to the bpf scheduler.
>
>
> > Introduce scx_bpf_task_proxy_cpu() and scx_bpf_task_proxy_cid() to
> > return the CPU or cid of the next mutex owner in the proxy chain.
> >
> > The owner relationship may change immediately after the query, so expose
> > the result only as a scheduling hint. Return a negative errno when no
> > valid proxy destination or cid mapping is available.
> >
> > Provide compatibility wrappers that return -EOPNOTSUPP when the kfuncs
> > are unavailable.
> >
> > Signed-off-by: Andrea Righi <arighi@xxxxxxxxxx>
> > ---
> > kernel/sched/core.c | 28 ++++++++++++++++
> > kernel/sched/ext/ext.c | 42 ++++++++++++++++++++++++
> > kernel/sched/ext/internal.h | 6 ++--
> > kernel/sched/sched.h | 2 ++
> > tools/sched_ext/include/scx/common.bpf.h | 2 ++
> > tools/sched_ext/include/scx/compat.bpf.h | 18 ++++++++++
> > 6 files changed, 96 insertions(+), 2 deletions(-)
> >
> > diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> > index 3d72f64ffe627..39e2689ea6c3b 100644
> > --- a/kernel/sched/core.c
> > +++ b/kernel/sched/core.c
> > @@ -7046,6 +7046,34 @@ find_proxy_task(struct rq *rq, struct task_struct *donor, struct rq_flags *rf)
> > return NULL;
> > }
> >
> > +int task_proxy_cpu(struct task_struct *p)
> > +{
> > + struct task_struct *owner;
> > + struct mutex *mutex;
> > +
> > + if (!sched_proxy_exec() || !READ_ONCE(p->is_blocked))
> > + return -ENOENT;
> > +
> > + guard(raw_spinlock_irqsave)(&p->blocked_lock);
> > +
> > + mutex = __get_task_blocked_on(p);
> > + if (!mutex)
> > + return -ENOENT;
> > +
> > + /*
> > + * @blocked_lock stabilizes @blocked_on and thus the mutex lifetime.
> > + * The owner is an atomic snapshot used only as a scheduling hint and
> > + * may change as soon as this function returns, so wait_lock is not
> > + * needed here.
> > + */
> > + owner = __mutex_owner(mutex);
> > + if (!owner)
> > + return -ENOENT;
> > + if (!READ_ONCE(owner->on_rq) || owner->se.sched_delayed)
> > + return -ENOENT;
> > +
> > + return task_cpu(owner);
> > +}
>
> So, I'm somewhat skeptical of this. You do disclaim in the commit log
> above that this is only a hint and it might change, but I fret it
> might be to a point it's not worth much as a hint.
>
> First: you're only looking at the immediate mutex owner, not the
> actual runnable owner of the full chain. So current could be on cpu1,
> the next owner on cpu2, but that task also just blocked on a mutex
> owned on cpu3. And the task on cpu2 may be about to migrate to cpu3
> (where current should ideally also proxy migrate to). So I'm not sure
> what such an ephemeral value might be worth. Again, maybe the
> optimization you have in mind is worth it, but it probably just needs
> some additional explanation in the commit message.
>
> Second: The blocked_lock does stabilize the mutex lifetime, but not
> the owner's. Once you've read the __mutex_owner(), without holding the
> mutex wait_lock, the mutex owner can be releasing the lock and then
> immediately exit, causing the owner dereferences that follow here to
> cause a UAF.

Yep, as mentioned in the previous email, let's get rid of this kfuncs completely
for now. In this way we can get rid of unnecessary complexity, fix these issues,
make the patch more self-consitent in sched_ext, it's a win-win. :)

Thanks,
-Andrea