Re: [PATCH v2] sched/deadline: Make dl-server nohz full aware

From: Ionut Nechita (Wind River)

Date: Mon Oct 05 2026 - 08:52:23 EST


Hi Juri,

Following up on the dl_servers_stop_all() hazard I mentioned earlier: you
said you had not been able to reproduce it, so here is a reliable
reproducer and the splats, in case it is useful before a v3 lands.

The setup is the one from my earlier mail (PREEMPT_RT, isolated nohz_full
CPU 2, fair_server defaults 50ms/1s), plus a churn of short-lived tasks on
the isolated core: a SCHED_FIFO hog and CFS tasks coexist, so the
fair_server is armed, while perf and a fork()/exit() loop keep coming and
going on that same CPU. That drives dl_server_start()/dl_server_stop()
through the "active contending" / "active non contending" running_bw
bandwidth-accounting state machine (with an inactive_timer in between) at
wakeup/dequeue rate.

When a stop finds the inactive_timer still armed, it subtracts running_bw
that the matching start never added back. running_bw underflows and trips
both WARNs:

- deadline.c:239 WARN_ON_ONCE(dl_rq->running_bw > old) in __sub_running_bw()
- deadline.c:425 WARN_ON(dl_se->dl_non_contending) in task_non_contending()

It fires from both the schedule path (pick_task_dl -> task_non_contending)
and the hrtimer path (hrtimer_interrupt -> inactive_task_timer). 15 splats
over ~4 minutes on this node; three representative ones (both WARN sites,
both paths) are below.

The change that triggers it, for reference, is an earlier version of my
own variant that stopped the server from the tick-stop paths (the version
I run now does not stop it there and is clean):

+static inline void sched_tick_stop_fair_server(struct rq *rq)
+{
+ if (!rq->rt.rt_nr_running || !rq->cfs.h_nr_queued)
+ dl_server_stop(&rq->fair_server);
+}
...
if (rq->rt.rr_nr_running) {
- if (rq->rt.rr_nr_running == 1)
+ if (rq->rt.rr_nr_running == 1) {
+ sched_tick_stop_fair_server(rq);
return true;
...
fifo_nr_running = rq->rt.rt_nr_running - rq->rt.rr_nr_running;
- if (fifo_nr_running)
+ if (fifo_nr_running) {
+ sched_tick_stop_fair_server(rq);
return true;

sched_can_stop_tick() runs on every add_nr_running()/sub_nr_running(), so
this stops a server that a wakeup a few microseconds later arms again.
This is structurally the same as the dl_servers_stop_all() call on the
RT-only path in your v2: a stop driven at enqueue/dequeue rate. To be
clear, I have not reproduced it with your v2 as posted, and it did not
fire in the measurement runs from my earlier mail - but it is the same
interleaving, so it seems worth confining the stop to the path that
actually stops the tick (as Peter also noted on the v2) rather than
running it on every sched_can_stop_tick() evaluation.

Representative splats (15 fired over ~4 minutes on the isolated core,
trimmed to three: both WARN sites and both the schedule and the hrtimer
path). Kernel 6.18.15-rt, PREEMPT_RT, isolated nohz_full CPU 2, Dell
PowerEdge R750. Short-lived tasks (perf, a fork/exit loop) churning on the
isolated core while a SCHED_FIFO hog and CFS tasks coexist, i.e. the
fair_server is armed and dl_server_start()/dl_server_stop() are being
cycled at wakeup/dequeue rate.

[ 3079.520930] WARNING: CPU: 2 PID: 410093 at kernel/sched/deadline.c:239 task_non_contending+0x242/0x390
[ 3079.521037] Call Trace:
[ 3079.521039] <TASK>
[ 3079.521042] pick_task_dl+0x50/0xb0
[ 3079.521046] __schedule+0x99a/0xfb0
[ 3079.521053] preempt_schedule_irq+0x2b/0x50
[ 3079.521055] asm_sysvec_irq_work+0x16/0x20
...
[ 3079.521081] </TASK>
[ 3079.521081] ---[ end trace 0000000000000000 ]---

[ 3087.363420] WARNING: CPU: 2 PID: 411091 at kernel/sched/deadline.c:425 task_non_contending+0x297/0x390
[ 3087.363525] Call Trace:
[ 3087.363526] <TASK>
[ 3087.363530] pick_task_dl+0x50/0xb0
[ 3087.363533] __schedule+0x99a/0xfb0
[ 3087.363538] schedule+0x23/0xd0
[ 3087.363541] do_wait+0x58/0x110
[ 3087.363544] kernel_wait4+0x9c/0x140
[ 3087.363548] __do_sys_wait4+0x36/0xb0
[ 3087.363571] do_syscall_64+0x7b/0xe30
[ 3087.363600] entry_SYSCALL_64_after_hwframe+0x76/0x7e
[ 3087.363614] </TASK>
[ 3087.363615] ---[ end trace 0000000000000000 ]---

[ 3087.363885] WARNING: CPU: 2 PID: 411092 at kernel/sched/deadline.c:239 inactive_task_timer+0x370/0x480
[ 3087.363948] Call Trace:
[ 3087.363949] <TASK>
[ 3087.363951] ? __pfx_inactive_task_timer+0x10/0x10
[ 3087.363953] __hrtimer_run_queues+0x145/0x280
[ 3087.363957] hrtimer_interrupt+0xf6/0x210
[ 3087.363960] __sysvec_apic_timer_interrupt+0x51/0xf0
[ 3087.363965] sysvec_apic_timer_interrupt+0x34/0x90
[ 3087.363970] asm_sysvec_apic_timer_interrupt+0x16/0x20
[ 3087.363980] </TASK>
[ 3087.363981] ---[ end trace 0000000000000000 ]---

Thanks,
Ionut