Re: [PATCH 2/2] sched/eevdf: Keep expired protection expired across reweighting

From: Christian Loehle

Date: Wed Sep 30 2026 - 11:34:20 EST


On 9/30/26 14:37, Kayra Cizmeci wrote:
> Hi Christian,
>
>> reweight_eevdf() rescales live slice protection, but leaves an expired
>> vprot unchanged when moving vruntime. A reweight can move vruntime behind
>> the old vprot. For example, reducing the weight of an entity with positive
>> lag can do so:
>>
>> before reweight: vprot <= vruntime
>> after reweight: vruntime < vprot
>>
>> protect_slice() consequently becomes true again, even though no fresh
>> protection was granted.
>
>> This also happens from task_tick_fair() after a queued HRTICK has
>> requested a reschedule. A four-task rt-app workload with 100 us, 1 ms,
>> 10 ms and 100 ms requests can then repick current despite a runnable,
>> eligible entity having an earlier deadline. The stale protection takes
>> precedence in pick_eevdf(). Cgroup weight changes also expose this with
>> HRTICK disabled.
>
>> Separating vprot from vlag allowed the expired absolute boundary to survive
>> the lag update and rescaling. Previously, those writes to vlag overwrote
>> the shared storage. Commit ff38424030f9 ("sched/eevdf: Update se->vprot in
>> reweight_entity()") subsequently handled live protection, but left the
>> expired case unchanged.
>
>> Keep an expired current entity's protection at its new vruntime:
>>
>> after reweight: vprot = vruntime
>>
>> This keeps protect_slice() false. Retain the existing rescaling for
>> protection that was still live and leave non-current entities alone.
>>
>> Fixes: 80390ead2080 ("sched/fair: Separate se->vlag from se->vprot")
>> Signed-off-by: Christian Loehle <christian.loehle@xxxxxxx>
>> ---
>> kernel/sched/fair.c | 3 +++
>> 1 file changed, 3 insertions(+)
>>
>> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
>> index 868c3911337a..d10eeea75f13 100644
>> --- a/kernel/sched/fair.c
>> +++ b/kernel/sched/fair.c
>> @@ -4951,6 +4951,9 @@ static void reweight_eevdf(struct cfs_rq *cfs_rq, struct sched_entity *se,
>> se->deadline += avruntime;
>> se->rel_deadline = 0;
>> se->vruntime = avruntime - se->vlag;
>> + /* Reweighting must not revive expired slice protection. */
>> + if (curr && !rel_vprot)
>> + se->vprot = se->vruntime;
>>
>> if (!curr)
>> __enqueue_entity(cfs_rq, se);
>
>
> Okay, so scene:
>
> we have a rq like this:
> +----+
> |root|
> +----+
> /\
> / \
> / \
> +------+ +------+
> |task_a| |task_b|
> +------+ +------+
>
> When, sched_change_end() activates for task_a, (Assuming it's Fair Class and weight is different than h_load.weight)
> enqueue_task_fair() calls reweight_eevdf() with on_rq = false so the block never works. Then we call place_entity() and vruntime goes back.
> On normal enqueue, this will be tolerated with vprot getting written over. But, sched_change_end() calls set_next_task()
> that calls set_next_task_fair() with SNT_NORMAL. On that case set_protect_slice() is not called. reweight_eevdf() is called
> again on that path, but because the first call did the job this one just returns early.
>
> I could be missing something tho. If I'm not, should we fold this into 2/2 or should I send a patch about this?
>

Thanks, looks legit to me and I did a quick rt-app + trace analysis to confirm.
I'll fold that in with you as a reporter if you don't mind.

I might start including the rt-app workloads in the cover-letter before I
lose track of them myself...