[tip: sched/urgent] sched: Account cgroup CPU time to the execution context
From: tip-bot2 for Hui Su
Date: Thu Sep 10 2026 - 05:46:53 EST
The following commit has been merged into the sched/urgent branch of tip:
Commit-ID: c23810313bdf6b02f39a1f2a1464c4b18bd39e31
Gitweb: https://git.kernel.org/tip/c23810313bdf6b02f39a1f2a1464c4b18bd39e31
Author: Hui Su <sh_def@xxxxxxx>
AuthorDate: Fri, 04 Sep 2026 11:47:07 +08:00
Committer: Peter Zijlstra <peterz@xxxxxxxxxxxxx>
CommitterDate: Thu, 10 Sep 2026 10:22:52 +02:00
sched: Account cgroup CPU time to the execution context
Proxy execution separates the scheduling context from the execution
context. Commit aa4f74dfd42b ("sched: Fix runtime accounting w/ split
exec & sched contexts") made per-task and thread-group runtime
accounting follow the task that actually executes, while cgroup CPU
usage is charged to the donor.
When the donor and execution task belong to different cgroups, this
makes a task's execution time count against a different cgroup from the
one the task belongs to.
Cgroup CPU usage should follow the execution context, matching the
per-task, thread-group, and cgroup user/system accounting. Keep
scheduling state associated with the donor, but charge cgroup CPU
usage to rq->curr.
A reproducer with the donor and execution task in separate cgroups
showed the execution task accumulating runtime while cgroup CPU usage
was charged to the donor's cgroup. With this change, the execution
task's cgroup accumulates the CPU usage instead. The same behavior was
verified with an RT donor and with legacy cpuacct accounting.
Fixes: aa4f74dfd42b ("sched: Fix runtime accounting w/ split exec & sched contexts")
Suggested-by: Tejun Heo <tj@xxxxxxxxxx>
Signed-off-by: Hui Su <sh_def@xxxxxxx>
Signed-off-by: Peter Zijlstra (Intel) <peterz@xxxxxxxxxxxxx>
Acked-by: Tejun Heo <tj@xxxxxxxxxx>
Acked-by: John Stultz <jstultz@xxxxxxxxxx>
Link: https://patch.msgid.link/20260904034707.268416-1-sh_def@xxxxxxx
---
kernel/sched/fair.c | 4 +---
1 file changed, 1 insertion(+), 3 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 944833e..7455a83 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1414,7 +1414,6 @@ static s64 update_se(struct rq *rq, struct sched_entity *se)
se->exec_start = now;
if (entity_is_task(se)) {
- struct task_struct *donor = task_of(se);
struct task_struct *running = rq->curr;
/*
* If se is a task, we account the time against the running
@@ -1427,8 +1426,7 @@ static s64 update_se(struct rq *rq, struct sched_entity *se)
account_group_exec_runtime(running, delta_exec);
account_mm_sched(rq, running, delta_exec);
- /* cgroup time is always accounted against the donor */
- cgroup_account_cputime(donor, delta_exec);
+ cgroup_account_cputime(running, delta_exec);
} else {
/* If not task, account the time against donor se */
se->sum_exec_runtime += delta_exec;