Re: [PATCH v2 4/4] sched/fair: Rework/fix task_h_load()

From: Peter Zijlstra

Date: Wed Sep 30 2026 - 04:46:01 EST


On Wed, Sep 30, 2026 at 06:46:38AM +0530, K Prateek Nayak wrote:
> Hello Kayra,
>
> On 9/29/2026 11:16 PM, Kayra Cizmeci wrote:
> > Hi Peter,
> >
> > (Fun Stuff):
> >
> >> @@ -15731,6 +15766,8 @@ static int __sched_group_set_shares(stru
> >> update_load_avg(cfs_rq, se, UPDATE_TG);
> >> update_cfs_group(se);
> >> }
> >> + for_each_sched_entity_bl(se, cfs_rq)
> >> + update_cfs_rq_h_load(group_cfs_rq(se), se, cfs_rq);
> >> rq_unlock_irqrestore(rq, &rf);
> >> }
> >>
> >
> > So, the code is this:
> > #define for_each_sched_entity(se, cfs_rq) \
> > for (struct sched_entity *_BL = NULL; \
> > (se) && ((cfs_rq) = cfs_rq_of(se), (cfs_rq)->backlink = _BL, true);\
> > (se) = (se)->parent, _BL = (se))
> >
> > #define for_each_sched_entity_bl(se, cfs_rq) \
> > for (; ((se) = (cfs_rq)->backlink); (cfs_rq) = group_cfs_rq(se))
> >
> >
> > (While writing this, a suitcase tried to assassinate me by falling from top of the closet,
> > what follows after this part may be the symptoms of my brain-damage.)

'Expect the unexpected!' :-)

Just to share the pictures I always noodle while doing this...

Given a cgroup hierarchy like:

Root
/ \
A B
| |
AA BA
| |
t1 t2

We have:

rq->cfs
/ \
cfs_rq-A - se-A se-B - cfs_rq-B
| |
cfs_rq-AA - se-AA se-BA - cfs_rq-BA
| |
se-1 - task_1 se-2 - task_2


Where:

- cfs_rq_of(se) - gives the cfs_rq se is enqueued on, iow. up one level
example: cfs_rq_of(se-AA) := cfs_rq-A

- group_cfs_rq(se) - gives the cfs_rq associated with the se, iow. sideways
example: group_cfs_rq(se-AA) := cfs_rq-AA

(I normally denote these as little arrows in the graph, but ASCII is a
little more rigid than pencil and paper.)


> > When we start as task_b, everything goes well. Both groups are updated.
> > On task_a too, only se_a is updated.
>
> I'm assuming the hierarchy is like the latter then if traversal from
> B updates A.
>
> >
> > But when we start as se_a root's backlink is NULL so we don't update anything.
> > While on se_b rq_a's backlink is NULL and update se_a but not ourselfes.
>
> So for that specific section you've highlighted from Peter's patch,
> in __sched_group_set_shares(), we first do a:
>
> for_each_sched_entity(se) {
> update_load_avg(cfs_rq_of(se), se, UPDATE_TG);
> update_cfs_group(se);
> }
>
> That sets up backlink going until se->parent whose group_cfs_rq() is
> the cfs_rq of tg_se(B) aka the cfs_rq just above the group whose shares
> were altered.
>
> Then we do:
>
> for_each_sched_entity_bl(se, cfs_rq)
> update_cfs_rq_h_load(group_cfs_rq(se), se, cfs_rq);
>
> Which updates the h_load all the way from the root until the cfs_rq of
> the cgroup we altered.
>
> In case of:
>
> root
> |
> A
> |
> B*
>
> *shares of cgroup is updated
>
> If we update shares of B (aka tg_se(B)), we update the h_load
> until the cfs_rq_of(tg_se(B)) which is till tg_cfs_rq(A).
>
> Now if you have:
>
> root
> |
> A
> / \
> *B C
> | \
> D E
>
> Yes,d you'll still update h_load for only A and you can have stale
> h_load for C, D, and E, and for all the tasks queued below them.
>
> Since full propagation is expensive, we do those propagation lazily
> when the task is picked, enqueued, or dequeued
>
> Note: We cannot propagate this up further because we have not yet done an
> update_load_avg() for the cfa_rq(s) in rest of the hierarchy. Next reweight
> will see the correct h_load starting from A and propagate it further when
> needed.
>
> Was that the problem you were talking about or did I totally confuse this
> with something else?

Right, so we only update the groups up from where we start. Ideally we'd
also update the whole affected subtree, but as Prateek says, that might
be expensive, and it will be updated on-demand later.