[PATCH 1/3] sched/fair: Drive NUMA and cache tick work from the execution context

From: Jemmy Wong

Date: Thu Oct 08 2026 - 09:31:08 EST


task_tick() is invoked with rq->donor. Under proxy execution the donor
is blocked on a mutex while rq->curr burns the CPU.

update_se() already charges sum_exec_runtime and LLC runtime to
rq->curr. The NUMA and cache tick hooks still used the donor, so
node_stamp + numa_scan_period was never crossed (donor runtime is
frozen) and TWA_RESUME task_work was queued on a blocked task.

Pass rq->curr for those hooks, matching psi_account_irqtime() and
wq_worker_tick(), and rename task_tick_fair()'s argument from curr to
donor so the scheduling context is not mistaken for rq->curr.

The lock holder need not be a fair task, so only run the hooks when
rq->curr is in fair_sched_class. This keeps NUMA scanning off non-fair
tasks and keeps task_tick_cache() consistent with account_mm_sched(),
which already ignores non-fair tasks.

Without CONFIG_SCHED_PROXY_EXEC the two rq members are a union.

Fixes: aa4f74dfd42b ("sched: Fix runtime accounting w/ split exec & sched contexts")
Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
Signed-off-by: Jemmy Wong <jemmywong512@xxxxxxxxx>
Assisted-by: LLM
---
kernel/sched/fair.c | 28 ++++++++++++++++++----------
1 file changed, 18 insertions(+), 10 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 57360f5cdde4..9dbc127a9ee1 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -15326,12 +15326,13 @@ static inline void task_tick_core(struct rq *rq, struct task_struct *curr) {}
*
* NOTE: This function can be called remotely by the tick offload that
* goes along full dynticks. Therefore no local assumption can be made
- * and everything must be accessed through the @rq and @curr passed in
+ * and everything must be accessed through the @rq and @donor passed in
* parameters.
*/
-static void task_tick_fair(struct rq *rq, struct task_struct *curr, int queued)
+static void task_tick_fair(struct rq *rq, struct task_struct *donor, int queued)
{
- struct sched_entity *se = &curr->se;
+ struct sched_entity *se = &donor->se;
+ struct task_struct *curr = rq->curr;

if (se->on_rq) {
unsigned long weight = NICE_0_LOAD;
@@ -15344,22 +15345,29 @@ static void task_tick_fair(struct rq *rq, struct task_struct *curr, int queued)
weight = __calc_prop_weight(cfs_rq, se, weight);
}

- se = &curr->se;
+ se = &donor->se;
reweight_eevdf(cfs_rq, se, weight, se->on_rq);
}

if (queued)
return;

- if (static_branch_unlikely(&sched_numa_balancing))
- task_tick_numa(rq, curr);
+ /*
+ * NUMA and cache work follow the execution context: update_se()
+ * charges runtime to rq->curr, and TWA_RESUME must run on that
+ * task. Under proxy-exec rq->curr may be of another class.
+ */
+ if (curr->sched_class == &fair_sched_class) {
+ if (static_branch_unlikely(&sched_numa_balancing))
+ task_tick_numa(rq, curr);

- task_tick_cache(rq, curr);
+ task_tick_cache(rq, curr);
+ }

- update_misfit_status(curr, rq);
- check_update_overutilized_status(task_rq(curr));
+ update_misfit_status(donor, rq);
+ check_update_overutilized_status(task_rq(donor));

- task_tick_core(rq, curr);
+ task_tick_core(rq, donor);
}

/*
--
2.54.0 (Apple Git-157)