Re: [PATCH v2 4/4] sched/fair: Rework/fix task_h_load()

From: Kayra Cizmeci

Date: Wed Sep 30 2026 - 07:53:05 EST


> > Hi Peter,
> >
> > I checked everything *again* to make sure;
> >
> > > - if (!se) {
> > > - cfs_rq->h_load = cfs_rq_load_avg(cfs_rq);
> > > - cfs_rq->last_h_load_update = now;
> > > - }
> > > + /*
> > > + * The above (forward) leaf_cfs_rq_list traversal will have done
> > > + * update_cfs_rq_load_avg() in a bottom-up fashion. Now iterate the
> > > + * list backwards, such that we're ensured to have visited every
> > > + * parent of the current group to update h_load in a top-down fashion.
> > > + */
> > > + list_for_each_entry_reverse(cfs_rq, &rq->leaf_cfs_rq_list, leaf_cfs_rq_list)
> > > + update_cfs_rq_h_load(cfs_rq, NULL, NULL);
> >
> >
> > > static unsigned long task_h_load(struct task_struct *p)
> > > {
> > > struct cfs_rq *cfs_rq = task_cfs_rq(p);
> > >
> > > - update_cfs_rq_h_load(cfs_rq);
> > > - return div64_ul(p->se.avg.load_avg * cfs_rq->h_load,
> > > + return div64_ul(p->se.avg.load_avg * READ_ONCE(cfs_rq->h_load),
> > > cfs_rq_load_avg(cfs_rq) + 1);
> > > }
> >
> > Is updating the root less often intentional?
> > Or am I missing something and we update the root more than the old code?

> You're referring to the removal of update_cfs_rq_h_load() here? That was
> the whole purpose of the patch. The callers of task_h_load() do not (in
> general) hold rq->lock, and thus update_cfs_rq_h_load() is unserialized
> and broken.

In the old code we were updating root when we go fully up and if it hadn't been
updated it already.

But now, that block is deleted. update_cfs_rq_h_load() is only called from this
other list_for_each_entry_reverse() block.

> Also, there's a typo:
>
> > + * Pelt uses apprixmate 'us' as ns/1024; and then uses time segments of 1024
> > + * 'us'. As a result each segment is in fact '1<<20' ns.
>
> 'apprixmate' :-).

> Some day I might learn to type :-)

My situation is worse than yours ;-)