Re: [PATCHSET sched_ext/for-7.4] sched_ext: Add NUMA balancing support
From: Vladimir Vdovin
Date: Tue Oct 06 2026 - 10:40:54 EST
Hi Andrea,
On Sun, Oct 04, 2026 at 09:27:06AM +0200, Andrea Righi wrote:
> Tasks running under a BPF scheduler are currently invisible to NUMA balancing:
> their address space is never scanned for hinting faults, so they never record
> NUMA fault statistics and never acquire a preferred node.
In case it's useful, I gave patches 1-3 a try on one of our KVM hosts,
backported to 6.18.54 and running our in-house sched_ext scheduler.
It's not exactly the series as posted: to avoid changing the scheduler,
the backport drops SCX_OPS_NUMA_BALANCING and drives the scan for all
sched_ext tasks under the sched_numa_balancing static key alone, and it
passes curr, as there is no donor there on 6.18. The fair.c parts are
the same.
As a controlled check, I used a stress-ng --vm worker touching 4G,
pinned to a CPU on node 0 and then moved to a CPU on node 1 with
taskset.
Without the patches, under the same scheduler, mm->numa_scan_seq stays
at 0, numa_preferred_nid stays at -1, numa_pte_updates in /proc/vmstat
doesn't move, and all 4G stay on node 0.
With the backport:
- mm->numa_scan_seq advances (7 to 12 over two minutes);
- numa_preferred_nid is 0 before the move and 1 about 20s after it;
- the whole 4G (~1M pages) is migrated to node 1 within ~20s;
- the task stays on the CPU it was pinned to.
On an otherwise idle host, the vCPU threads of a running guest now get
task_numa_work() calls as well. On a host running the same scheduler
without the patches, a kprobe on task_numa_work() sees no calls at all
in 30s, as task_tick_numa() is only reached from task_tick_fair().
Thanks for working on this,
Vladimir