Re: [PATCH 2/2] sched/eevdf: Keep expired protection expired across reweighting

From: Kayra Cizmeci

Date: Wed Sep 30 2026 - 10:23:38 EST


Hi Christian,

> reweight_eevdf() rescales live slice protection, but leaves an expired
> vprot unchanged when moving vruntime. A reweight can move vruntime behind
> the old vprot. For example, reducing the weight of an entity with positive
> lag can do so:
>
> before reweight: vprot <= vruntime
> after reweight: vruntime < vprot
>
> protect_slice() consequently becomes true again, even though no fresh
> protection was granted.

> This also happens from task_tick_fair() after a queued HRTICK has
> requested a reschedule. A four-task rt-app workload with 100 us, 1 ms,
> 10 ms and 100 ms requests can then repick current despite a runnable,
> eligible entity having an earlier deadline. The stale protection takes
> precedence in pick_eevdf(). Cgroup weight changes also expose this with
> HRTICK disabled.

> Separating vprot from vlag allowed the expired absolute boundary to survive
> the lag update and rescaling. Previously, those writes to vlag overwrote
> the shared storage. Commit ff38424030f9 ("sched/eevdf: Update se->vprot in
> reweight_entity()") subsequently handled live protection, but left the
> expired case unchanged.

> Keep an expired current entity's protection at its new vruntime:
>
> after reweight: vprot = vruntime
>
> This keeps protect_slice() false. Retain the existing rescaling for
> protection that was still live and leave non-current entities alone.
>
> Fixes: 80390ead2080 ("sched/fair: Separate se->vlag from se->vprot")
> Signed-off-by: Christian Loehle <christian.loehle@xxxxxxx>
> ---
> kernel/sched/fair.c | 3 +++
> 1 file changed, 3 insertions(+)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 868c3911337a..d10eeea75f13 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -4951,6 +4951,9 @@ static void reweight_eevdf(struct cfs_rq *cfs_rq, struct sched_entity *se,
> se->deadline += avruntime;
> se->rel_deadline = 0;
> se->vruntime = avruntime - se->vlag;
> + /* Reweighting must not revive expired slice protection. */
> + if (curr && !rel_vprot)
> + se->vprot = se->vruntime;
>
> if (!curr)
> __enqueue_entity(cfs_rq, se);


Okay, so scene:

we have a rq like this:
+----+
|root|
+----+
/\
/ \
/ \
+------+ +------+
|task_a| |task_b|
+------+ +------+

When, sched_change_end() activates for task_a, (Assuming it's Fair Class and weight is different than h_load.weight)
enqueue_task_fair() calls reweight_eevdf() with on_rq = false so the block never works. Then we call place_entity() and vruntime goes back.
On normal enqueue, this will be tolerated with vprot getting written over. But, sched_change_end() calls set_next_task()
that calls set_next_task_fair() with SNT_NORMAL. On that case set_protect_slice() is not called. reweight_eevdf() is called
again on that path, but because the first call did the job this one just returns early.

I could be missing something tho. If I'm not, should we fold this into 2/2 or should I send a patch about this?


Thanks,
Kayra