Re: [PATCH v2 4/4] sched/fair: Rework/fix task_h_load()

From: Kayra Cizmeci

Date: Thu Oct 01 2026 - 17:17:17 EST


> > > Hi Peter,
> > >
> > > I checked everything *again* to make sure;
> > >
> > > > - if (!se) {
> > > > - cfs_rq->h_load = cfs_rq_load_avg(cfs_rq);
> > > > - cfs_rq->last_h_load_update = now;
> > > > - }
> > > > + /*
> > > > + * The above (forward) leaf_cfs_rq_list traversal will have done
> > > > + * update_cfs_rq_load_avg() in a bottom-up fashion. Now iterate the
> > > > + * list backwards, such that we're ensured to have visited every
> > > > + * parent of the current group to update h_load in a top-down fashion.
> > > > + */
> > > > + list_for_each_entry_reverse(cfs_rq, &rq->leaf_cfs_rq_list, leaf_cfs_rq_list)
> > > > + update_cfs_rq_h_load(cfs_rq, NULL, NULL);
> > >
> > >
> > > > static unsigned long task_h_load(struct task_struct *p)
> > > > {
> > > > struct cfs_rq *cfs_rq = task_cfs_rq(p);
> > > >
> > > > - update_cfs_rq_h_load(cfs_rq);
> > > > - return div64_ul(p->se.avg.load_avg * cfs_rq->h_load,
> > > > + return div64_ul(p->se.avg.load_avg * READ_ONCE(cfs_rq->h_load),
> > > > cfs_rq_load_avg(cfs_rq) + 1);
> > > > }
> > >
> > > Is updating the root less often intentional?
> > > Or am I missing something and we update the root more than the old code?

> > You're referring to the removal of update_cfs_rq_h_load() here? That was
> > the whole purpose of the patch. The callers of task_h_load() do not (in
> > general) hold rq->lock, and thus update_cfs_rq_h_load() is unserialized
> > and broken.

> In the old code we were updating root when we go fully up and if it hadn't been
> updated it already.

> But now, that block is deleted. update_cfs_rq_h_load() is only called from this
> other list_for_each_entry_reverse() block.

Ah.. Why I'm like this? (Maybe it's because of the brain-damage :>)

I meant that the only place that the root was getting updated were this list_for_each_entry_reverse()
block.

Old code was giving maximum 1 jiffy obsolescence (But it was broken, as you say it)
while the new one is less broken but doesn't gives any obsolescence clarity.

Nevermind, I wanna talk from the code that should'a be easier. (BTW, my git tree is mess, everything is everywhere
it's like a warzone. Hope I'm not the only one messy with these things...)


diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 50583fde20e1..aa9a54817dc6 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -5773,8 +5773,10 @@ static inline void __update_cfs_rq_h_load(struct cfs_rq *cfs_rq,
if (!p_cfs_rq)
p_cfs_rq = cfs_rq_of(se);

- load = p_cfs_rq->h_load;
- load = div64_ul(load * se->avg.load_avg,
+ if (p_cfs_rq == &p_cfs_rq->rq->cfs)
+ load = se->avg.load_avg;
+ else
+ load = div64_ul(p_cfs_rq->h_load * se->avg.load_avg,
p_cfs_rq->avg.load_avg + 1);
}

@@ -11617,6 +11619,9 @@ static unsigned long task_h_load(struct task_struct *p)
{
struct cfs_rq *cfs_rq = task_cfs_rq(p);

+ if (cfs_rq == &cfs_rq->rq->cfs)
+ return p->se.avg.load_avg;
+
return div64_ul(p->se.avg.load_avg * READ_ONCE(cfs_rq->h_load),
cfs_rq_load_avg(cfs_rq) + 1);
}

Yeah, well I was thinking about something like this.
Maybe I'm misunderstanding something tho :/,
I done this in the middle of the night, untested.
I should sleep. Ah..

> Also, there's a typo:
>
> > + * Pelt uses apprixmate 'us' as ns/1024; and then uses time segments of 1024
> > + * 'us'. As a result each segment is in fact '1<<20' ns.
>
> 'apprixmate' :-).

> Some day I might learn to type :-)

> My situation is worse than yours ;-)

Anyway,
Kayra