Re: [PATCH v8 0/2] hung_task: Improve warning budget handling and task reporting
From: Aaron Tomlin
Date: Thu Aug 06 2026 - 10:05:50 EST
On Thu, Aug 06, 2026 at 10:05:10AM +0800, Lance Yang wrote:
>
>
> On 2026/8/5 22:16, Aaron Tomlin wrote:
> > On Wed, Aug 05, 2026 at 10:13:05AM +0800, Lance Yang wrote:
> > >
> > >
> > > On 2026/8/5 07:05, Andrew Morton wrote:
> > > > On Tue, 4 Aug 2026 16:20:48 -0400 Aaron Tomlin <atomlin@xxxxxxxxxxx> wrote:
> > > >
> > > > > The hung_task watchdog detects tasks stuck in TASK_UNINTERRUPTIBLE (D)
> > > > > state for longer than CONFIG_DEFAULT_HUNG_TASK_TIMEOUT seconds. To prevent
> > > > > log spam during system spikes, sysctl_hung_task_warnings enforces a budget
> > > > > on the number of logged warnings.
> > > > >
> > > > > However, the current implementation has two major limitations:
> > > > >
> > > > > 1. Permanent exhaustion of warning budget
> > > > >
> > > > > sysctl_hung_task_warnings is decremented directly when printing
> > > > > warnings. Once this budget hits zero, no further warnings are
> > > > > reported until an administrator manually updates the sysctl value or
> > > > > reboots the system. Consequently, a single temporary hang episode
> > > > > permanently blinds the kernel watchdog to any subsequent hung tasks
> > > > > after system recovery.
> > > > >
> > > > > 2. Total log suppression when budget is exhausted
> > > > >
> > > > > Once the warning budget reaches zero, hung_task_info() completely
> > > > > suppresses all output, including the basic single-line alert. While
> > > > > suppressing verbose stack dumps and lock debugging is desirable to
> > > > > prevent dmesg flooding, hiding basic task alerts leaves
> > > > > administrators entirely unaware that tasks are hanging.
> > > > >
> > > > > This patch series resolves both limitations by decoupling the configured
> > > > > warning budget from the runtime warning counter, automatically resetting
> > > > > the budget when the system recovers, and keeping basic single-line hung
> > > > > task alerts visible.
> > > >
> > > > Thanks. A couple of concerns from AI review:
> > > > https://sashiko.dev/#/patchset/20260804202050.262427-1-atomlin@xxxxxxxxxxx
> > >
> > > I'm not quite sure what the cleanest way to handle these is yet, but
> > > both points look fair.
> > >
> > > 1) Concurrent writes to hung_task_warnings can race and leave
> > > hung_task_warnings_printed out of sync with it.
> > >
> > > 2) The unconditional pr_err() is also no longer bounded by
> > > hung_task_warnings. With lots of hung tasks, every scan can flood
> > > the log and console with one line per task. Maybe rate-limit those
> > > messages or cap them per scan.
> > >
> > > > Apologies if these were considered during review of previous
> > > > iterations.
> >
> > Hi Andrew, Lance,
> >
> > Yes. However, I feel the first one is of a lesser concern. For instance,
> > consider the following race scenario, when two threads write to the sysctl
> > concurrently:
> >
> > - Thread A writes value 10, 'writes sysctl_hung_task_warnings = 10'
> > - Thread B writes value 20, 'writes sysctl_hung_task_warnings = 20'
> > - Thread B executes 'hung_task_warnings_printed =
> > sysctl_hung_task_warnings' (i.e., sets 20)
> >
> > - Thread A resumes and executes 'hung_task_warnings_printed =
> > sysctl_hung_task_warnings' using its _cached_ register value 10
> >
> > The result, sysctl_hung_task_warnings holds 20, but
> > hung_task_warnings_printed holds 10.
> >
> > I suspect the severity is low since concurrent sysctl writes are likely
> > rare—restricted to CAP_SYS_ADMIN. Finally, if de-synchronisation occurs,
> > the system automatically self-heals as soon as a watchdog check finds zero
> > hung tasks (this_round_count == 0), resetting hung_task_warnings_printed =
> > sysctl_hung_task_warnings.
> >
> > However, I would rather not leave the data race unresolved. How about using
> > READ_ONCE() and WRITE_ONCE()? I think multi-variable transactional
> > atomicity is unnecessary:
>
> Doesn't close the race. A can read 10, B can finish both updates
> with 20, then A writes 10 back. Still ends up 20/10.
>
> READ_ONCE()/WRITE_ONCE() don't serialize anything here ...
Yes; READ_ONCE() and WRITE_ONCE() guarantee single-copy load/store
atomicity to prevent compiler optimisations. However, I agree, they do not
establish a critical section or serialise multi-step operations across
variables 'sysctl_hung_task_warnings' and 'hung_task_warnings_printed'.
How about the following?
diff --git a/kernel/hung_task.c b/kernel/hung_task.c
index 6ebb3a87ac65..a977dc280262 100644
--- a/kernel/hung_task.c
+++ b/kernel/hung_task.c
@@ -237,6 +237,8 @@ static inline void debug_show_blocker(struct task_struct *task, unsigned long ti
static void hung_task_info(struct task_struct *t, unsigned long timeout,
unsigned long this_round_count)
{
+ int warnings;
+
trace_sched_process_hang(t);
if (sysctl_hung_task_panic && this_round_count >= sysctl_hung_task_panic) {
@@ -254,9 +256,11 @@ static void hung_task_info(struct task_struct *t, unsigned long timeout,
* CONFIG_DEFAULT_HUNG_TASK_TIMEOUT. Therefore, complain
* accordingly with full details if the budget is not exhausted.
*/
- if (hung_task_warnings_printed || hung_task_call_panic) {
- if (hung_task_warnings_printed > 0)
- hung_task_warnings_printed--;
+ warnings = READ_ONCE(hung_task_warnings_printed);
+
+ if (warnings || hung_task_call_panic) {
+ if (warnings > 0)
+ WRITE_ONCE(hung_task_warnings_printed, warnings - 1);
pr_err(" %s %s %.*s\n",
print_tainted(), init_utsname()->release,
(int)strcspn(init_utsname()->version, " "),
@@ -268,7 +272,7 @@ static void hung_task_info(struct task_struct *t, unsigned long timeout,
sched_show_task(t);
debug_show_blocker(t, timeout);
- if (!hung_task_warnings_printed)
+ if (!READ_ONCE(hung_task_warnings_printed))
pr_info("Future hung task reports won't print details about each process, see sysctl kernel.hung_task_warnings\n");
}
@@ -308,7 +312,7 @@ static void check_hung_uninterruptible_tasks(unsigned long timeout)
unsigned long last_break = jiffies;
struct task_struct *g, *t;
unsigned long this_round_count;
- int need_warning = hung_task_warnings_printed;
+ int need_warning = READ_ONCE(hung_task_warnings_printed);
unsigned long si_mask = hung_task_si_mask;
/*
@@ -345,7 +349,7 @@ static void check_hung_uninterruptible_tasks(unsigned long timeout)
rcu_read_unlock();
if (!this_round_count) {
- hung_task_warnings_printed = sysctl_hung_task_warnings;
+ WRITE_ONCE(hung_task_warnings_printed, READ_ONCE(sysctl_hung_task_warnings));
return;
}
@@ -431,20 +435,21 @@ static int proc_dohung_task_timeout_secs(const struct ctl_table *table, int writ
return ret;
}
+static DEFINE_MUTEX(hung_task_sysctl_mutex);
+
static int proc_dohung_task_warnings(const struct ctl_table *table, int write,
void *buffer,
size_t *lenp, loff_t *ppos)
{
int ret;
+ mutex_lock(&hung_task_sysctl_mutex);
ret = proc_dointvec_minmax(table, write, buffer, lenp, ppos);
+ if (!ret && write)
+ WRITE_ONCE(hung_task_warnings_printed, READ_ONCE(sysctl_hung_task_warnings));
+ mutex_unlock(&hung_task_sysctl_mutex);
- if (ret || !write)
- return ret;
-
- hung_task_warnings_printed = sysctl_hung_task_warnings;
-
- return 0;
+ return ret;
}
/*
> > diff --git a/kernel/hung_task.c b/kernel/hung_task.c
> > index 4e1fb0db79d1..123456789abc 100644
> > --- a/kernel/hung_task.c
> > +++ b/kernel/hung_task.c
> > @@ -348,7 +349,7 @@ static void check_hung_uninterruptible_tasks(unsigned long timeout)
> >
> > if (!this_round_count) {
> > - hung_task_warnings_printed = sysctl_hung_task_warnings;
> > + WRITE_ONCE(hung_task_warnings_printed, READ_ONCE(sysctl_hung_task_warnings));
> > return;
> > }
> >
> > @@ -429,14 +430,14 @@ static int proc_dohung_task_warnings(const struct ctl_table *table, int write,
> > void *buffer,
> > size_t *lenp, loff_t *ppos)
> > {
> > int ret;
> >
> > ret = proc_dointvec_minmax(table, write, buffer, lenp, ppos);
> >
> > if (ret || !write)
> > return ret;
> >
> > - hung_task_warnings_printed = sysctl_hung_task_warnings;
> > + WRITE_ONCE(hung_task_warnings_printed, READ_ONCE(sysctl_hung_task_warnings));
> >
> > return 0;
> > }
> >
> > For the second issue, this is very serious. We could move the per-task
> > blocked message back inside the budget check and when the budget is
> > exhausted, emit a single aggregate summary line at the end of
> > check_hung_uninterruptible_tasks(). However, this is not ideal.
>
> Yeah, aggregate summary works. I'd drop timeout, though. It's just
> scan threshold, not actual blocked time, and summary no longer refers
> to any one task. Maybe just:
>
> pr_info("khungtaskd: %lu hung tasks detected (warning budget exhausted)\n",
> this_round_count);
Acknowledged.
Kind regards,
--
Aaron Tomlin