[tip: sched/urgent] sched/fair: Use update_curr_eevdf() for the remaining root cfs_rq callers
From: tip-bot2 for Zhan Xusheng
Date: Wed Sep 02 2026 - 03:26:15 EST
The following commit has been merged into the sched/urgent branch of tip:
Commit-ID: 1719d035a6fa90b7467b6daf45a573f5180013b2
Gitweb: https://git.kernel.org/tip/1719d035a6fa90b7467b6daf45a573f5180013b2
Author: Zhan Xusheng <zhanxusheng@xxxxxxxxxx>
AuthorDate: Sat, 22 Aug 2026 18:59:30 +08:00
Committer: Peter Zijlstra <peterz@xxxxxxxxxxxxx>
CommitterDate: Wed, 02 Sep 2026 09:17:49 +02:00
sched/fair: Use update_curr_eevdf() for the remaining root cfs_rq callers
pick_task_fair() and yield_task_fair() call update_curr(&rq->cfs) to bring
curr up to date before they look at the eevdf state. With cgroups that
does not happen: update_curr() reads ->h_curr, which on the root cfs_rq is
the top level group entity, and returns at the !entity_is_task() check
before touching vruntime. Both then read ->curr, so the guard and the
update disagree about which entity they mean.
Counting how often ->h_curr and ->curr differ at pick_task_fair(), on one
CPU for 10s with three busy tasks and one 200us-periodic task:
all tasks in the root cgroup 43321 calls, 0 no-ops
busy tasks in G0, periodic in G1 45211 calls, 45193 no-ops
Whether that matters depends on what precedes the pick. Since
commit 68e37487810a ("sched/fair: Fix flat hierarchy") the tick and
enqueue/dequeue all update curr correctly, so on the normal reschedule
path only the microseconds between those and the pick are missing, and I
could not measure a latency difference there. Three paths have nothing
before them on that rq though:
- pick_task() on the sibling rqs of a core under core scheduling
(kernel/sched/core.c), which updates that rq's clock first for
exactly this reason
- fair_server_pick_task()
- yield_task_fair(), where the stale value feeds the entity_eligible()
test that guards forfeiting the remaining vruntime
There curr can be a full tick behind, as it was before that commit.
No new behaviour for the entity being updated: without cgroups ->h_curr
is already the task, so these two call sites already run the full
update_curr() including update_deadline(), dl_server_update() and the
resched_curr_lazy() at the end. This makes the cgroup case do the same.
Fixes: 85570f10a4c6 ("sched/eevdf: Move to a single runqueue")
Signed-off-by: Zhan Xusheng <zhanxusheng@xxxxxxxxxx>
Signed-off-by: Peter Zijlstra (Intel) <peterz@xxxxxxxxxxxxx>
Reviewed-by: Vincent Guittot <vincent.guittot@xxxxxxxxxx>
Link: https://patch.msgid.link/20260822105930.2352761-1-zhanxusheng1024@xxxxxxxxx
---
kernel/sched/fair.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 8dff370..5d47de5 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -10057,7 +10057,7 @@ again:
/* Might not have done put_prev_entity() */
if (cfs_rq->curr && cfs_rq->curr->on_rq)
- update_curr(cfs_rq);
+ update_curr_eevdf(cfs_rq);
se = pick_next_entity(rq, true);
if (!se)
@@ -10160,7 +10160,7 @@ static void yield_task_fair(struct rq *rq)
/*
* Update run-time statistics of the 'current'.
*/
- update_curr(cfs_rq);
+ update_curr_eevdf(cfs_rq);
/*
* Tell update_rq_clock() that we've just updated,
* so we don't do microscopic update in schedule()