Re: [PATCH] Differentiate scenarios when watchdog is closed
From: Guenter Roeck
Date: Thu Aug 27 2026 - 14:12:06 EST
On 8/27/26 10:09, chaithco@xxxxxxxxxx wrote:
On Thu, Aug 27 2026 at 09:01:47 AM -07:00:00, Guenter Roeck <linux@xxxxxxxxxxxx> wrote:
Presently
[...]
Also, the subject should start with the subsystem name ("watchdog:")
[...]
deliberately
Thank you for catching these! Please accept my apologies. I can fix those up in the next submission.
[...] Also, while technically userspace may close the
watchdog deliberately while it is running, that is not what happens
on a regular basis.
This is actually what initiated a bug report at https://bugzilla.redhat.com/show_bug.cgi?id=1991285 it turns out systemd explicitly does this to help ensure a system shutting down actually eventually goes down even if the shutdown process hits some snags. It does this on every shutdown. Given the prevalence of systemd, this is a regular occurrence. The end result is that, when using iTCO, it shows an error on every shutdown when systemd is in use as init.
If you want to make a change, I would suggest to add an error message
into watchdog_stop() to report an error if the stop callback returns
an error. That would distinguish 2/3 without making functional changes.
Thank you! So something like this?
if (wdd->ops->stop) {
clear_bit(WDOG_HW_RUNNING, &wdd->status);
err = wdd->ops->stop(wdd);
+ if (err < 0)
+ pr_info("watchdog%d: closed while still enabled!\n");
More like
pr_err(""watchdog%d: Failed to stop watchdog: %pe\n", wdd->id, ERR_PTR(err));
since this would be a real error.
The "watchdog%d: watchdog did not stop!" message will then follow
(unconditionally).
trace_watchdog_stop(wdd, err);
} else {
set_bit(WDOG_HW_RUNNING, &wdd->status);
While responding to this, an additional thought occurred to me; given the primary reason a user would see this is because systemd is shutting down a system, it may be more worth while to have systemd log something about closing the watchdog without disarming it to at least explain a pr_crit kernel log line. Otherwise, it just looks like "something bad happened" with watchdog. I am additionally unsure of what would be best to go in watchdog_stop that helps differentiate intentional closing of the watchdog without disabling vs malicious/accidental closing. The intent would lie within the entity closing the watchdog; "closed while still enabled!" still seems like "something bad happened" with info on if it was intentional or not.
Problem is that we don't know if "something bad happened". The same message
will be seen if the watchdog daemon was killed or crashed. We can not just
assume that closing the watchdog device was intentional.
Guenter