Re: [PATCH v9 1/2] hung_task: Reset warning budget when problem gets resolved

From: Lance Yang

Date: Fri Aug 28 2026 - 05:36:50 EST




On 2026/8/28 17:05, Petr Mladek wrote:
On Thu 2026-08-27 23:30:01, Lance Yang wrote:
On Wed, Aug 26, 2026 at 01:17:47PM +0200, Petr Mladek wrote:
@@ -314,6 +317,11 @@ static void check_hung_uninterruptible_tasks(unsigned long timeout)
if (test_taint(TAINT_DIE) || did_panic)
return;
+ if (atomic_xchg(&reset_hung_task_warnings, 0))

I would use here atomic_xchg_acquire(). It serializes the ordering
of reset_hung_task_warnings vs sysctl_hung_task_warnings.
It would make it symetric with the barrier in the sysctl handler.

Yep, _acquire is enough here. Plain atomic_xchg() is already fully
ordered, though, so this looks like making the intent clearer rather
than fixing the ordering :)

Yes, my intention was to make the ordering more clear and symmetric.

Yep, that makes sense. atomic_xchg_acquire() is a better fit here :)

The original code worked because the barrier was even stronger.

The old-value return already makes plain atomic_xchg() fully ordered :)

ORDERING (see memory-barriers.txt)
--------

The rule of thumb:
...
- RMW operations that have a return value are fully ordered;
...
Except of course when a successful operation has an explicit ordering
like:

{}_relaxed: unordered
{}_acquire: the R of the RMW (or atomic_read) is an ACQUIRE
{}_release: the W of the RMW (or atomic_set) is a RELEASE


+ hung_task_warnings_printed =
+ READ_ONCE(sysctl_hung_task_warnings);

This would work only when "sysctl_hung_task_warnings"
is updated using WRITE_ONCE(). But it seems that this
is not the case. My understading is that it is updated by:

Wait, I think proc_dointvec_minmax() already handles this.

You are right.

For proc_dointvec_minmax(), the converter is:

int proc_dointvec_minmax(const struct ctl_table *table, int dir,
void *buffer, size_t *lenp, loff_t *ppos)
{
return do_proc_dointvec(table, dir, buffer, lenp, ppos,
do_proc_int_conv_minmax);
}

Here, i is table->data, while lval is local:

I have missed this.

No worries at all. This was easy to miss in that call chain ...


static int do_proc_dointvec(const struct ctl_table *table, int dir,
void *buffer, size_t *lenp, loff_t *ppos,
int (*conv)(bool *negp, unsigned long *u_ptr, int *k_ptr,
int dir, const struct ctl_table *table))
{
...
i = (int *) table->data;
vleft = table->maxlen / sizeof(*i);
...
for (; left && vleft--; i++, first=0) {
unsigned long lval;
bool neg;

if (SYSCTL_USER_TO_KERN(dir)) {
proc_skip_spaces(&p, &left);

if (!left)
break;
err = proc_get_long(&p, &left, &lval, &neg,
proc_wspace_sep,
sizeof(proc_wspace_sep), NULL);

I have missed that proc_get_long() assigns the value to the local
variable @lval.

if (err)
break;
if (conv(&neg, &lval, i, 1, table)) {
err = -EINVAL;
break;

The real asigment to table->data is done here. And I agree that it
goes down to proc_int_conv() which does WRITE_ONCE().

So, we are on the safe side and do _not_ need the proxy table.

Agreed.


Now, I am not sure whether we need v10. It might be worth it.
AFAIK, Andrew has not taken this patchset yet...

I think we do. Definitely :) And if Andrew hasn't picked it up
yet, even better. There's still time to fold the changes in :)


I am sorry for complications.

No need to apologize at all, Petr. I really appreciate you taking
another careful look!

Cheers, Lance