[PATCH] sched/fair: Avoid overflow in place_entity()
From: Hui Su
Date: Mon Sep 28 2026 - 09:45:48 EST
place_entity() inflates an entity's virtual lag with
lag * (load + weight) / load
Commit 4823725d9d1d ("sched/fair: Increase weight bits for avg_vruntime")
replaced the scaled-down weights used in this calculation with
avg_vruntime_weight(), removing the arithmetic headroom provided by
scale_load_down().
The multiplication can overflow s64 before the division even when the
final quotient is representable. A task's hierarchical effective weight
can be as small as 2 after avg_vruntime_weight() scaling under group
scheduling, while entity_lag() bounds |vlag| by
calc_delta_fair(cfs_rq_max_slice(cfs_rq) + TICK_NSEC, se)
For example, with HZ=1000, NICE_0_LOAD=1048576, and the minimum scaled
hierarchical weight scale_load(2)=2048, this bound is 891289600. With 114
maximum-weight entities, the runqueue load is 10361604096. The
intermediate product is:
891289600 * (10361604096 + 2)
= 9235189971864780800
> S64_MAX
The signed overflow corrupts the entity's placement.
Rewrite
lag * (load + weight) / load
as
lag + lag * weight / load
The two expressions are equivalent, but the latter avoids multiplying the
virtual lag by the total runqueue weight. entity_lag() bounds vlag against
the entity's effective weight, so the remaining lag * weight product stays
bounded. rescale_entity() preserves this scale across h_load changes and
already relies on the same product when rescaling vlag.
Keep the original zero-load warning semantics by multiplying lag by weight
after replacing the divisor with one.
Fixes: 4823725d9d1d ("sched/fair: Increase weight bits for avg_vruntime")
Signed-off-by: Hui Su <sh_def@xxxxxxx>
---
Testing:
- PASS: UBSAN positive/negative overflow cases, zero-load/control paths,
64-bit landmarks NICE_0_LOAD=1048576 and scale_load(88761)=90891264,
unscaled 32-bit landmarks 1024/88761, and 100000 random cases checked
against a __int128 oracle.
- PASS: Full x86_64 kernel build; kernel/sched/fair.o builds for i386
and arm64.
kernel/sched/fair.c | 15 ++++++++++++---
1 file changed, 12 insertions(+), 3 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 7455a83a6a99..a013bfba576b 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -6268,10 +6268,19 @@ place_entity(struct cfs_rq *cfs_rq, struct sched_entity *se, int flags)
load += avg_vruntime_weight(cfs_rq, curr->h_load.weight);
weight = avg_vruntime_weight(cfs_rq, se->h_load.weight);
- lag *= load + weight;
- if (WARN_ON_ONCE(!load))
+ if (WARN_ON_ONCE(!load)) {
load = 1;
- lag = div64_long(lag, load);
+ lag *= weight;
+ } else {
+ /*
+ * Avoid overflowing lag * (load + weight) by distributing
+ * the division:
+ *
+ * lag * (load + weight) / load
+ * = lag + lag * weight / load
+ */
+ lag += div64_long(lag * weight, load);
+ }
/*
* A heavy entity (relative to the tree) will pull the
base-commit: 6812ce4e4379ffc99c52401ec28f0d7ffbc36206