Re: [PATCH 3/5] sched_ext: Scan NUMA hinting faults for opted-in BPF schedulers
From: Andrea Righi
Date: Mon Oct 05 2026 - 17:37:48 EST
Hi Vladimir,
On Mon, Oct 05, 2026 at 02:30:50PM +0300, Vladimir Vdovin wrote:
> Hi Andrea,
>
> Thanks for putting this together so quickly, and for picking up the RFC.
>
> On Sun, Oct 04, 2026 at 09:27:09AM +0200, Andrea Righi wrote:
> > + if (!queued && (sch->ops.flags & SCX_OPS_NUMA_BALANCING) &&
> > + static_branch_unlikely(&sched_numa_balancing))
> > + task_tick_numa(rq, donor);
>
> One question about proxy execution, which I may well be missing context
> on. Hui Su's tick series moves the fair NUMA tick to the execution
> context, with the reasoning that task_tick_numa() works on the mm and
> NUMA work state of the task that is actually running:
>
> https://lore.kernel.org/r/20260909092901.2989564-3-sh_def@xxxxxxx
>
> Here the scan is driven from the donor. I understand that in your proxy
> execution integration SCX runtime is accounted to the donor, so maybe
> this is intentional to keep the pacing consistent, but then the scan
> would cover the donor's address space rather than the one being
> accessed. Is the donor the intended choice here, or should this follow
> rq->curr once SCHED_PROXY_EXEC no longer depends on !SCHED_CLASS_EXT?
>
> Thanks,
> Vladimir
Yes, that's a good point, thanks for brining this up. The donor here is used
just for consistency with the current task_tick_fair() implementation.
A small clarification on the accounting: under proxy exec, sched_ext charges the
slice to the donor (update_curr_scx() decrements donor->scx.slice), but the task
runtime, p->se.sum_exec_runtime, is charged to rq->curr by update_se(). Since
task_tick_numa() paces the scan on sum_exec_runtime, passing the donor means its
scan does not progress while it's blocked, which is fine, and the owner's scan
is delayed until it gets ticks on its own. So nothing is scanned on the wrong
mm: task_tick_numa(rq, p) checks whether p is due for a scan and, if so, queues
p's scan as task work, which runs later in p's own context and p's own mm. The
owner never scans the donor's memory or the other way around. Essentially, the
only effect of passing the donor is timing: during the proxy window the donor's
scan may get queued, while the owner's scan waits until the owner is schedued on
its own behalf.
That said, I agree the scan should follow rq->curr and Hui's series gives the
right hook to do it. But I'd rather not diverge from fair for now, so the idea
would be to keep the donor here and switch sched_ext together with fair when
Hui's series lands.
Thanks,
-Andrea