[PATCH 2/5] sched/numa: Leave the placement of a BPF-scheduled task to its scheduler

From: Andrea Righi

Date: Sun Oct 04 2026 - 03:29:49 EST


A task under a BPF scheduler can take NUMA hinting faults via
task_numa_fault() and then reach numa_migrate_preferred(), which may
migrate the task to its preferred node through task_numa_migrate().

The migration of such a task should always be driven by its BPF
scheduler, which owns the placement of its tasks.

Keep collecting the statistics, which is what makes
p->numa_preferred_nid worth reading, but skip the migration for a task
owned by a BPF scheduler. The scheduler can read the preferred node and
decide for itself whether moving the task is worth the cost.

Reported-by: Vladimir Vdovin <deliran@xxxxxxxxxx>
Link: https://lore.kernel.org/r/20261002124559.10367-1-deliran@xxxxxxxxxx
Signed-off-by: Andrea Righi <arighi@xxxxxxxxxx>
---
kernel/sched/fair.c | 8 ++++++++
1 file changed, 8 insertions(+)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index e2d52faacdb7a..019c1331545bf 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -3446,6 +3446,14 @@ static void numa_migrate_preferred(struct task_struct *p)
if (task_node(p) == p->numa_preferred_nid)
return;

+ /*
+ * A task under a BPF scheduler is placed by that scheduler. Keep the
+ * statistics coming, which is what makes p->numa_preferred_nid worth
+ * reading, but leave the placement alone.
+ */
+ if (task_on_scx(p))
+ return;
+
/* Otherwise, try migrate to a CPU on the preferred node */
task_numa_migrate(p);
}
--
2.55.0