[PATCH v2 4/5] sched/eevdf: Keep expired protection expired across reweighting
From: Christian Loehle
Date: Thu Oct 01 2026 - 09:48:22 EST
reweight_eevdf() rescales live slice protection, but leaves an expired
vprot unchanged when moving vruntime. A reweight can move vruntime behind
the old vprot. For example, reducing the weight of an entity with positive
lag can do so:
before reweight: vprot <= vruntime
after reweight: vruntime < vprot
protect_slice() consequently becomes true again, even though no fresh
protection was granted.
This can also happen in the HRTICK callback. entity_tick() can expire
protection and request a reschedule, after which task_tick_fair() calls
reweight_eevdf() before schedule() gets a chance to select another entity.
The reweight can move vruntime behind the stale vprot, making
protect_slice() true again before the pending selection. A four-task rt-app
workload with 100 us, 1 ms, 10 ms and 100 ms requests can then repick
current despite a runnable, eligible entity having an earlier deadline.
Cgroup weight changes also expose this with HRTICK disabled.
sched_change_begin() exposes another path by temporarily dequeuing current.
enqueue_task_fair() then reweights it with on_rq clear before
place_entity() moves vruntime. Remember whether protection was expired
before placement so the later SNT_NORMAL restore cannot revive it.
Separating vprot from vlag allowed the expired absolute boundary to
survive the lag update and rescaling. Previously, those writes to vlag
overwrote the shared storage. Commit ff38424030f9 ("sched/eevdf: Update
se->vprot in reweight_entity()") subsequently handled live protection,
but left the expired case unchanged.
Keep an expired current entity's protection at its new vruntime:
after reweight: vprot = vruntime
This keeps protect_slice() false. Retain the existing rescaling for
protection that was still live and leave non-current entities alone.
Fixes: 80390ead2080 ("sched/fair: Separate se->vlag from se->vprot")
Reported-by: Kayra Cizmeci <kayracizmeci@xxxxxxxxx>
Link: https://lore.kernel.org/lkml/20260930133716.214471-1-kayracizmeci@xxxxxxxxx/
Signed-off-by: Christian Loehle <christian.loehle@xxxxxxx>
---
kernel/sched/fair.c | 9 ++++++++-
1 file changed, 8 insertions(+), 1 deletion(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 15fa967273f6..bffc560bc7b8 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -4941,6 +4941,9 @@ static void reweight_eevdf(struct cfs_rq *cfs_rq, struct sched_entity *se,
se->deadline += avruntime;
se->rel_deadline = 0;
se->vruntime = avruntime - se->vlag;
+ /* Reweighting must not revive expired slice protection. */
+ if (curr && !rel_vprot)
+ se->vprot = se->vruntime;
if (!curr)
__enqueue_entity(cfs_rq, se);
@@ -8204,7 +8207,7 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
struct sched_entity *se = &p->se;
struct cfs_rq *cfs_rq = &rq->cfs;
unsigned long weight;
- bool curr;
+ bool curr, expired = false;
if (task_is_throttled(p) && enqueue_throttled_task(p))
return;
@@ -8237,6 +8240,8 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
* XXX comment on the curr thing
*/
curr = (cfs_rq->curr == se);
+ if (!curr && task_current_donor(rq, p))
+ expired = !protect_slice(se);
if (curr)
place_entity(cfs_rq, se, flags);
@@ -8248,6 +8253,8 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
if (!curr) {
reweight_eevdf(cfs_rq, se, weight, false);
place_entity(cfs_rq, se, flags | ENQUEUE_QUEUED);
+ if (expired)
+ se->vprot = se->vruntime;
__enqueue_entity(cfs_rq, se);
}
--
2.34.1