Re: [PATCH 0/3] watchdog: omap: Preserve OMAP watchdog across boot to detect faulty boots
From: Diogo Ivo
Date: Mon Sep 14 2026 - 04:39:01 EST
On 9/11/26 7:25 PM, Guenter Roeck wrote:
On 9/11/26 08:17, Diogo Ivo wrote:
Hi Guenter,I did. I just wonder if the effort is worth the pain / cost.
On 9/11/26 4:29 PM, Guenter Roeck wrote:
On 9/11/26 02:17, Diogo Ivo wrote:
Allow the OMAP watchdog to survive being stopped on kernel initializationPlease address the issues reported by Sashiko, or explain why they don't apply.
so it can detect a faulty boot in cases where the bootloader leaves it
running and the watchdog driver picks it up during kernel init.
- Patch 1 removes a duplicate omap_wdt_start() call left behind by
cd004d8299f1 ("watchdog: Fix OMAP watchdog early handling"). This is
unrelated to the main goal of the series and can be picked up
independently.
- Patch 2 adds support for reading the watchdog boot status. Probe now
checks whether the watchdog is already running and takes it over instead
of blindly stopping it based on early_enable alone. This introduces a
regression possibility, explained in detail in the patch's message.
- Patch 3 marks the OMAP4 watchdog node ti,no-reset-on-init so the
ti-sysc driver stops resetting it. A detailed explaination of why is
also provided in the commit message of the patch.
This series has been tested on a platform based on the VAR-SOM-OM44 from
Variscite, running a TI OMAP4460 SoC.
I have just replied to the Sashiko reviews but I'm not sure if you got
the replies as Sashiko did not include your e-mail in its review. If you
did not receive them please let me know and I can resend them. In
any case if you could give your opinion on the comments I left on the
patches about regressions that would be great as I think after the
Sashiko points are addressed that is the main blocker for this series.
Is there an actual use case ? Is the problem you are trying to solve
a real problem, or a theoretic one ? For example, the patches impose
a hard boot delay of more than 30 ms in omap_wdt_is_running().
Even though that could be optimized (there is no reason to wait
that long; the value could change a microsecond after the first read),
it is nevertheless a mandatory boot delay.
Yes, I stumbled upon this problem while trying to do exactly what is on
the commit messages, so having a system that does A/B updates and uses
the watchdog to detect if the boot succeeded or failed. The current boot
delay is actually (1000000 / 32768) * 10 = 305us, so quite a bit smaller
than the 30ms you mention, even though from the Sashiko comments I need
to take into account the prescaler, which in the worst case brings the
value to the 30ms you mention, but if I adjust it I can bring it down to
3ms in the worst case of a 128 prescaler. However, for a system with
prescaler=1 the delay can be brought down to 30us.
Another concern is the impact and potential side effects of setting
ti,no-reset-on-init (and the possible boot loop cause by it due to the odd
30-second init delay). After this change, a running watchdog is no longer
stopped. What happens on systems which do not load the watchdog at all
(for example because the driver was not configured) ? Will that also cause
a boot loop on such systems ?
Here with ti,no-reset-on-init the ti-sysc driver will shutdown the
watchdog driver after 30 seconds, which can indeed cause problems. When
I sent the patch I thought the timeout would be 3 seconds, which
considerable reduces the possibility of a bootloop. I will look into it
to see if it makes sense to change the timeout value and reduce the
possibility of the bootloop. My initial assumption of the 3 seconds was
what actually made me not add ti,no-idle-on-reset since that would
completely block the watchdog stopping from ti-sysc.
This is just a couple of problems introduced by this series. You better have
a very good reason for it to warrant having to deal with the potential fallout.
The reason is that in these systems the watchdog is not behaving as one
would expect: if the bootloader leaves the watchdog running it is a
reasonable expectation that unless the kernel or userspace services it
the system should shutdown. With this series I have tried to achieve
this while minimizing the risk of regressions, and from my point of
view the only sore point is indeed the 30s timeout in ti-sysc. If you
agree with my reasoning and think this is worth pursuing let me know and
I will change the 30s timeout for v2. In the meantime patch 1 is
completely regression free and can be picked up!
Thanks,
Diogo
Thanks,
Guenter