Re: [PATCH v9 1/2] hung_task: Reset warning budget when problem gets resolved

From: Petr Mladek

Date: Wed Aug 26 2026 - 07:18:28 EST


Hi,

first, I am sorry for so late review. I had vacation, many
things accumulated, ...

On Fri 2026-08-14 09:57:17, Aaron Tomlin wrote:
> The sysctl hung_task_warnings currently holds both the configured warning
> limit and the remaining budget. Each detailed report decrements the
> sysctl, so once it reaches zero, the configured limit is lost and cannot
> be restored automatically.
>
> Keep sysctl hung_task_warnings unchanged and track the remaining budget
> in hung_task_warnings_printed. Reset the runtime budget via an atomic flag
> when a watchdog check sees no hung tasks or when userspace writes a new
> sysctl value.
>
> --- a/kernel/hung_task.c
> +++ b/kernel/hung_task.c
> @@ -59,6 +59,9 @@ static unsigned long __read_mostly sysctl_hung_task_check_interval_secs;
>
> static int __read_mostly sysctl_hung_task_warnings = 10;
>
> +static int hung_task_warnings_printed = 10;

Nit: The name of the variable is a bit misleading in this final
version. It does not longer count the number of printed
messages.

A better name might be "hung_task_warnings_budget" or so.

It would be nice to change it if we need another version.
And I am afraid that we would need it, see below.

If we change it then I would also add comments explaining
the difference between the two values, something like:

<proposal>
/*
* Limit the number of printed hung tasks to prevent printing
* the same or similar backtraces repeatedly.
*/
static int __read_mostly sysctl_hung_task_warnings = 10;
/*
* The number of hung tasks which still can be reported.
* The budget gets restored to the original limit when
* the previous stall is resolved.
*/
static int hung_task_warnings_budget = 10;
</proposal>

> +static atomic_t reset_hung_task_warnings = ATOMIC_INIT(0);
> +
> static int __read_mostly did_panic;
> static bool hung_task_call_panic;
>
> @@ -245,11 +248,11 @@ static void hung_task_info(struct task_struct *t, unsigned long timeout,
> /*
> * The given task did not get scheduled for more than
> * CONFIG_DEFAULT_HUNG_TASK_TIMEOUT. Therefore, complain
> - * accordingly
> + * accordingly with full details if the budget is not exhausted.
> */
> - if (sysctl_hung_task_warnings || hung_task_call_panic) {
> - if (sysctl_hung_task_warnings > 0)
> - sysctl_hung_task_warnings--;
> + if (hung_task_warnings_printed || hung_task_call_panic) {
> + if (hung_task_warnings_printed > 0)
> + hung_task_warnings_printed--;
> pr_err("INFO: task %s:%d blocked%s for more than %ld seconds.\n",
> t->comm, t->pid, t->in_iowait ? " in I/O wait" : "",
> (jiffies - t->last_switch_time) / HZ);
> @@ -264,7 +267,7 @@ static void hung_task_info(struct task_struct *t, unsigned long timeout,
> sched_show_task(t);
> debug_show_blocker(t, timeout);
>
> - if (!sysctl_hung_task_warnings)
> + if (!hung_task_warnings_printed)
> pr_info("Future hung task reports are suppressed, see sysctl kernel.hung_task_warnings\n");
> }
>
> @@ -304,7 +307,7 @@ static void check_hung_uninterruptible_tasks(unsigned long timeout)
> unsigned long last_break = jiffies;
> struct task_struct *g, *t;
> unsigned long this_round_count;
> - int need_warning = sysctl_hung_task_warnings;
> + int need_warning;
> unsigned long si_mask = hung_task_si_mask;
>
> /*
> @@ -314,6 +317,11 @@ static void check_hung_uninterruptible_tasks(unsigned long timeout)
> if (test_taint(TAINT_DIE) || did_panic)
> return;
>
> + if (atomic_xchg(&reset_hung_task_warnings, 0))

I would use here atomic_xchg_acquire(). It serializes the ordering
of reset_hung_task_warnings vs sysctl_hung_task_warnings.
It would make it symetric with the barrier in the sysctl handler.

> + hung_task_warnings_printed =
> + READ_ONCE(sysctl_hung_task_warnings);

This would work only when "sysctl_hung_task_warnings"
is updated using WRITE_ONCE(). But it seems that this
is not the case. My understading is that it is updated by:

+ proc_dointvec_minmax()
+ do_proc_vec()
+ proc_get_long()
+ strtoul_lenient()

which does a plain assigment:

static int strtoul_lenient(const char *cp, char **endp, unsigned int base,
unsigned long *res)
{
[...]
*res = (unsigned long)result;
[...]
}

It can be solved by using temporary variable in proc_dointvec_minmax().
We have a custom proc_dohung_task_warnings() handler anyway.
See below.

> + need_warning = hung_task_warnings_printed;
> +
> this_round_count = 0;
> rcu_read_lock();
> for_each_process_thread(g, t) {
> @@ -340,8 +348,11 @@ static void check_hung_uninterruptible_tasks(unsigned long timeout)
> unlock:
> rcu_read_unlock();
>
> - if (!this_round_count)
> + if (!this_round_count) {
> + hung_task_warnings_printed =
> + READ_ONCE(sysctl_hung_task_warnings);
> return;
> + }
>
> if (need_warning || hung_task_call_panic) {
> si_mask |= SYS_INFO_LOCKS;
> @@ -425,6 +436,19 @@ static int proc_dohung_task_timeout_secs(const struct ctl_table *table, int writ
> return ret;
> }
>
> +static int proc_dohung_task_warnings(const struct ctl_table *table, int write,
> + void *buffer,
> + size_t *lenp, loff_t *ppos)
> +{
> + int ret;
> +
> + ret = proc_dointvec_minmax(table, write, buffer, lenp, ppos);
> + if (!ret && write)
> + atomic_set_release(&reset_hung_task_warnings, 1);
> +
> + return ret;
> +}

We should use WRITE_ONCE() when updating proc_dohung_task_warnings.
So, we need similar trick with proxy_table like in
proc_dohung_task_detect_count. Something like, on top of this patch:

--- a/kernel/hung_task.c
+++ b/kernel/hung_task.c
@@ -444,13 +444,26 @@ static int proc_dohung_task_warnings(const struct ctl_table *table, int write,
void *buffer,
size_t *lenp, loff_t *ppos)
{
+ struct ctl_table proxy_table;
+ int warnings;
int ret;

- ret = proc_dointvec_minmax(table, write, buffer, lenp, ppos);
- if (!ret && write)
- atomic_set_release(&reset_hung_task_warnings, 1);
+ proxy_table = *table;
+ proxy_table.data = &warnings;

- return ret;
+ if (SYSCTL_KERN_TO_USER(write))
+ warnings = READ_ONCE(sysctl_hung_task_warnings);
+
+ ret = proc_dointvec_minmax(&proxy_table, write, buffer, lenp, ppos);
+ if (ret < 0)
+ return ret;
+
+ if (SYSCTL_USER_TO_KERN(write)) {
+ WRITE_ONCE(sysctl_hung_task_warnings, warnings);
+ atomic_set_release(&reset_hung_task_warnings, 1);
+ }
+
+ return 0;
}

/*

Otherwise, it looks good to me.

Best Regards,
Petr