I have shipped enough embedded devices to know that no matter how carefully you review code, your firmware will eventually stop doing what it is supposed to do. It might be an I2C sensor that holds the bus low, a priority inversion that starves a critical task, or a single bit flip from EMI in an industrial cabinet. Without a recovery mechanism, that device sits in the field doing nothing until someone power-cycles it. The watchdog timer is the simplest, cheapest insurance policy we have for embedded reliability, and in my experience, it is also one of the most misunderstood peripherals. A watchdog does not fix bugs. It guarantees that when your system inevitably faults, it recovers on its own. For any product that runs unattended, that self-recovery behavior is non-negotiable.
Why Your Firmware Will Eventually Hang Without a Watchdog Timer
When I bring up a new board, the watchdog is one of the first peripherals I enable, even before the application code is stable. The reason is simple: embedded systems fail in ways that desktop software does not. We deal with tight memory, no operating system to kill a runaway process, direct hardware interaction, and environments that are electrically hostile. A loop that waits for a flag that never sets, a driver that blocks forever on a peripheral that went offline, or heap fragmentation that causes an allocation to never return can all freeze your system without triggering a hard fault.
A watchdog timer is essentially a hardware countdown timer that will reset the microcontroller if it is not periodically refreshed, or "kicked." If your main loop or RTOS is running correctly, it kicks the dog before the timer expires. If your firmware hangs, the kick never happens and the hardware forces a reset. That reset puts you back into a known good state.
The kinds of faults watchdogs actually catch
In my own projects, the most common watchdog saves have not been from wild pointer bugs. They have been from subtle system-level issues. I once had a LoRa-based sensor node that would freeze once every three weeks because the radio driver entered an infinite wait for a DIO pin that occasionally glitched. No hard fault, no exception, just a dead task. The watchdog reset the node in under two seconds and the device rejoined the network. Without it, the node would have required a 50km drive to reset manually.
Other scenarios I see repeatedly include: a high-priority task stuck in a while loop polling a sensor that never responds, a deadlock between two tasks competing for a SPI bus, and stack overflow that corrupts the task control block so the scheduler can no longer switch tasks. If you are working with an RTOS, remember that a deadlock will not crash the CPU; it will just make it look like everything has stopped. This is precisely where careful RTOS design matters, and understanding how tasks interact is critical. I often point junior engineers to RTOS Synchronization: Mutexes, Semaphores and Event Groups because improper use of those primitives is a leading cause of the hangs that watchdogs end up rescuing.
What a watchdog cannot do for you
A watchdog is a fault recovery mechanism, not a fault prevention mechanism. It will not tell you why you reset, it will not save RAM contents unless you design it to, and it will not help if you kick it from the wrong place. I have seen firmware where the watchdog was kicked from a high-priority timer interrupt that kept running even after the main application tasks were completely deadlocked. The system never reset, which was worse than having no watchdog at all because it created a false sense of embedded reliability. The kick must represent real proof that your application is healthy.
Independent Watchdogs, Window Watchdogs, and Software Supervisors Compared
Most modern microcontrollers offer more than one type of watchdog. Choosing the right one depends on what kind of failures you want to detect. On STM32, MSP430, and most Zephyr-supported SoCs, you will encounter two hardware types: the Independent Watchdog (IWDG) and the Window Watchdog (WWDG). Some systems also implement a software watchdog supervisor on top of the hardware.
The Independent Watchdog runs from its own dedicated low-speed clock, typically an internal RC oscillator around 32 kHz or 40 kHz. Because it is independent from the main system clock, it will still reset the MCU even if your main crystal fails or the system clock is misconfigured. It has a relatively coarse timeout, usually from a few milliseconds to tens of seconds, and you can kick it at any time before it expires.
The Window Watchdog runs from the main system clock (APB1 on STM32, for example) and is much more precise. It introduces a window: you must kick it not too late, but also not too early. If you kick it before the window opens, it resets you. This catches a different class of bugs, like a runaway loop that kicks the watchdog too frequently, or an interrupt storm that keeps kicking even though the main tasks are dead.
In my experience, IWDG is mandatory for basic fault recovery, while WWDG is excellent for safety-critical tasks where timing correctness matters. For complex RTOS systems, I usually combine the hardware IWDG with a software supervisor that monitors individual tasks.
| Feature | Independent Watchdog (IWDG) | Window Watchdog (WWDG) | Software Task Supervisor |
|---|---|---|---|
| Clock Source | Dedicated low-speed RC oscillator | System clock (APB / bus clock) | RTOS tick, no dedicated hardware |
| Timeout Range | ~1 ms to >30 s, coarse resolution | ~100 µs to ~50 ms, fine resolution | Configurable per task, ms to seconds |
| Early Kick Detection | No, kick anytime before expiry | Yes, reset if kicked too early | Yes, detects tasks kicking too often or not enough |
| Survives System Clock Failure | Yes | No | No |
| Best Use Case | General hang and deadlock recovery | Timing-critical and safety functions | Detecting single stuck task in RTOS |
If you are building a product on Zephyr, the watchdog API abstracts this nicely. The Zephyr Project Documentation describes the wdt API that handles both IWDG and WWDG under a common driver model, which helps when you need to port between hardware.
How Watchdog Timers Fit Into RTOS Task Scheduling and the Idle Hook
On bare metal, kicking the watchdog from your main super-loop is straightforward. With an RTOS, the question becomes: which task should kick the dog? The answer is none of them directly. If a single low-priority task kicks the watchdog, a higher-priority task can be dead and you will never know. If every task kicks it independently, you lose the ability to detect partial failures.
The pattern that has worked reliably for me in FreeRTOS and Zephyr is a centralized supervisor. One high-priority supervisor task is the only entity allowed to kick the hardware watchdog. All other important tasks must periodically check in with the supervisor to prove they are alive. If any task fails to check in within its expected window, the supervisor intentionally stops kicking the hardware and lets the reset happen.
Why kicking from the idle task is usually a mistake
Many FreeRTOS examples show kicking the watchdog from the Idle hook (vApplicationIdleHook). I avoid this. The idle task runs only when no other task is ready to run. If a high-priority task enters an infinite loop and never blocks, the idle task never executes, so the watchdog will reset - that part works. But if a medium-priority task is stuck waiting for a queue that never arrives, the idle task may still run and keep kicking, masking the failure. The idle task tells you the CPU has free cycles, not that your application is healthy. The details of how priorities, preemption and time slicing affect this are covered well in FreeRTOS Task Scheduling: Priorities, Preemption and Time Slicing, and I recommend reviewing it before you decide where your watchdog kick lives.
A reliable place to kick: a dedicated supervisor task
Instead, create a supervisor task at a high priority that wakes up every 100-200 ms, checks flags or counters updated by other tasks, and only then踢s the hardware watchdog. This gives you per-task monitoring with a single hardware timer. The supervisor should run more frequently than the watchdog timeout, but not so frequently that it adds jitter. I typically set the hardware timeout to 1-2 seconds and have the supervisor run every 250 ms.
/* STM32 IWDG bare-metal setup - runs on LSI ~32kHz */
#include "stm32f4xx.h"
void iwdg_init(uint32_t timeout_ms) {
// Enable write access to IWDG registers
IWDG->KR = 0x5555;
// Set prescaler to 32 -> 32kHz / 32 = 1kHz tick (1ms per count)
IWDG->PR = 3;
// Calculate reload value: RLR = timeout_ms * (LSI / prescaler) / 1000
// For 1000ms timeout: RLR = 1000
uint16_t reload = (timeout_ms * 1);
if (reload > 0xFFF) reload = 0xFFF;
IWDG->RLR = reload;
// Start watchdog
IWDG->KR = 0xCCCC;
// Kick it once to load reload value
IWDG->KR = 0xAAAA;
}
void iwdg_kick(void) {
IWDG->KR = 0xAAAA;
}
Calculating Timeout Windows: Too Short Means Spurious Resets, Too Long Means Bricked Devices
Selection of the timeout period is where I see the most field failures. Too short, and normal system latency causes false resets during flash writes, network joins, or sensor calibration. Too long, and your device remains unresponsive for an unacceptable time after a fault, which hurts embedded reliability metrics and can violate safety requirements.
I size the timeout by measuring the worst-case execution time of the longest atomic operation that cannot be split. For example, erasing a flash page on an nRF52 can block for 90 ms. An MQTT publish with TLS handshake on a poor cellular link might block a task for 2 seconds. If your watchdog is set to 500 ms, those normal operations will trigger it.
My practical formula for timeout selection
First, identify the longest period where you intentionally cannot kick. Add 50% margin, then add the supervisor period. In my experience, a 1.5 to 2-second IWDG timeout works for most sensor nodes. For gateway devices that perform firmware updates, I sometimes go to 4-5 seconds, but I also temporarily increase the timeout or use a windowed approach during the update to avoid spurious resets.
For window watchdogs, you also need to calculate the lower bound. The window is typically defined by a down-counter. On STM32 WWDG, the counter starts at 0x7F and counts down to 0x40. You can only kick when the counter is below the window value. This prevents a stuck high-priority interrupt from kicking too fast. When I use WWDG for a motor control loop, I set the window so the kick must occur within the last 30% of the period, which proves the control loop executed at the correct rate.
Handling long blocking operations without disabling the watchdog
Never disable the watchdog to get around a long operation. Instead, either kick it from within the long operation if you can prove forward progress, or split the operation into smaller chunks that let the supervisor run. For flash operations, I use a state machine that erases one page, checks in with the supervisor, then continues. The FreeRTOS Documentation provides guidance on breaking blocking calls into non-blocking patterns that work well with this approach. Disabling the watchdog during OTA updates is a common cause of bricked fleets when the update itself hangs.
/* Zephyr window watchdog example - must kick inside window */
#include <zephyr/drivers/watchdog.h>
const struct device *wdt = DEVICE_DT_GET(DT_ALIAS(watchdog0));
void wdt_init_window(void) {
struct wdt_timeout_cfg cfg = {
.window = {
.min = 400, /* cannot kick before 400ms */
.max = 800, /* must kick before 800ms */
},
.flags = WDT_FLAG_RESET_SOC,
.callback = NULL,
};
wdt_install_timeout(wdt, &cfg);
wdt_setup(wdt, 0);
}
void control_loop(void) {
while (1) {
run_pid_update(); /* takes ~5ms, must be periodic */
k_sleep(K_MSEC(500));
/* Kick only here - if loop runs too fast or too slow, reset occurs */
wdt_feed(wdt, 0);
}
}
Building a Multi-Task Watchdog Manager That Actually Catches Silent Failures
The most effective pattern I have used in production RTOS firmware is a software watchdog manager that sits between tasks and the hardware watchdog. Each critical task increments a counter or sets a flag when it completes a full cycle of useful work. The manager verifies all counters are progressing, then kicks the hardware. A task that is blocked, deadlocked, or starved stops checking in, the manager detects it, and the system resets.
This catches silent failures that a hardware watchdog alone would miss. I had a case where a sensor task was waiting forever on a semaphore that was never given due to an I2C NACK handling bug. The task was alive from the scheduler's point of view, it was just blocked. The hardware watchdog kept getting kicked by other tasks, so the device appeared healthy but reported stale data for days. After adding task-level monitoring, that bug caused a reset within 3 seconds and the error was logged.
What counts as a valid check-in
A good check-in is not just that the task is running, but that it did useful work. Incrementing a counter at the top of the task loop proves nothing. I require tasks to check in only after they have completed a full iteration: read sensor, process, send to queue successfully, then check in. For a communications task, I check in only after a successful MQTT ping or queue drain. This way, a task that is spinning but not making progress is still detected.
You should also consider dependencies on memory. A task that hangs inside a heap allocation due to fragmentation will not check in, but the root cause is memory. If you are seeing frequent supervisor resets, review your heap strategy and whether you should be using static allocation. The article on RTOS Memory Management: Static Allocation, Pools and Heap Strategies has a solid breakdown of how allocation failures manifest as task hangs.
/* FreeRTOS multi-task watchdog supervisor - only this task kicks hardware */
#include "FreeRTOS.h"
#include "task.h"
#include "timers.h"
#define NUM_MONITORED_TASKS 3
#define SUPERVISOR_PERIOD_MS 250
#define TASK_TIMEOUT_MS 1000
static volatile uint32_t task_heartbeat[NUM_MONITORED_TASKS] = {0};
static volatile uint32_t task_last_seen[NUM_MONITORED_TASKS] = {0};
/* Call this from each monitored task after completing useful work */
void watchdog_checkin(uint8_t task_id) {
task_heartbeat[task_id]++;
}
void vWatchdogSupervisorTask(void *pvParameters) {
TickType_t last_wake = xTaskGetTickCount();
for (;;) {
vTaskDelayUntil(&last_wake, pdMS_TO_TICKS(SUPERVISOR_PERIOD_MS));
bool all_ok = true;
TickType_t now = xTaskGetTickCount();
for (int i = 0; i < NUM_MONITORED_TASKS; i++) {
if (task_heartbeat[i] == task_last_seen[i]) {
/* No progress since last check */
if ((now - task_last_seen[i]) * portTICK_PERIOD_MS > TASK_TIMEOUT_MS) {
all_ok = false;
/* Optional: log which task failed before reset */
// log_error("WDT: task %d starved", i);
}
} else {
task_last_seen[i] = task_heartbeat[i];
}
}
if (all_ok) {
iwdg_kick(); /* Only kick hardware if all tasks healthy */
}
/* If all_ok == false, we intentionally starve the hardware WDT */
}
}
/* Example monitored task */
void vSensorTask(void *pvParameters) {
for (;;) {
if (read_sensor_and_queue() == 0) {
watchdog_checkin(0); /* Check in only on success */
}
vTaskDelay(pdMS_TO_TICKS(500));
}
}
Key details: the supervisor task should be one of the highest priority tasks so it is not starved itself. I also add a startup grace period where the supervisor does not enforce checks for the first 5 seconds after boot, to allow all tasks to initialize. After that grace period, every task must check in regularly. If you use this pattern, test it by deliberately blocking one task with a vTaskDelay(10000) and confirming the system resets within the expected window.
Tracing the Cause of a Watchdog Reset When You Have No Debugger Attached
A watchdog reset that leaves no trace is almost as bad as no watchdog. In the lab you have JTAG, but in the field you need to store enough information to diagnose the cause after the reset. Every watchdog implementation I ship stores the reset reason and, if possible, the state of each task at the time of failure.
Most MCUs have a reset status register that distinguishes watchdog reset from power-on, brown-out, or software reset. On STM32, check RCC_CSR, on nRF52 check NRF_POWER->RESETREAS, on ESP32 check rtc_get_reset_reason(). Log this early in boot before you clear the flags. I reserve a small section of no-init RAM or RTC backup registers that survive a watchdog reset to store the last task check-in values and program counter.
Preserving forensic data across the reset
Place a structure in a .noinit section so it is not zeroed by startup code. Before the supervisor decides to let the watchdog expire, it can write which task timed out into that RAM. After reset, your boot code can read it and send it to your cloud or log it to flash. This turns an anonymous reset into an actionable bug report.
For deeper analysis, use trace tools during development. When I debug intermittent watchdog resets, I run with a hardware trace probe capturing task switches and interrupts. Techniques described in Real-Time Debugging: Trace Tools, Logic Analyzers and JTAG Techniques are invaluable for correlating a watchdog reset with the exact sequence that led to it. In one case, SEGGER SystemView showed that a low-priority logging task was being starved for 1.8 seconds because a higher-priority task held a mutex too long - exactly the issue our supervisor was designed to catch.
Testing your watchdog like a fault-injection harness
Do not ship a watchdog you have not tested under fault conditions. My checklist includes: 1) Halt the CPU with a debugger and confirm reset within timeout+margin, 2) Block each monitored task individually with an infinite loop or semaphore take and confirm the supervisor lets the hardware expire, 3) Test early-kick detection for window watchdogs by kicking in a tight loop, 4) Measure current during reset to ensure peripherals are properly re-initialized. I also run a soak test where I intentionally inject I2C failures and heap pressure for 72 hours to see if the watchdog fires at all. If it never fires during soak, I may have made the timeout too long or the check-ins too permissive.
Frequently Asked Questions
Should I use the watchdog during development and debugging?
Yes, but configure it so it does not interfere with halting. Most MCUs have a DBGMCU or similar register that pauses the watchdog when the core is halted by a debugger. On STM32 set DBGMCU_APB1_FZ2_IWDG_STOP, on nRF52 set WDT_CONFIG with pause on halt. In my experience, leaving the watchdog running during development catches hangs early, but you must pause it on break or you will reset while stepping through code. For initial bring-up, I use a longer timeout like 4 seconds and shorten it to the production value once timing is validated.
What is the difference between kicking the watchdog from an interrupt versus a task?
Kicking from an interrupt is risky because interrupts continue to run even when tasks are deadlocked. A high-priority timer interrupt that kicks the watchdog will mask most RTOS failures. I only allow kicks from a task context, specifically the supervisor task. If you must kick from an interrupt, make that interrupt also verify a flag set by a task, so the kick still proves task-level health. The only exception I make is using the window watchdog from a precise timer interrupt when the goal is to prove interrupt timing correctness.
Can a watchdog help with stack overflow or memory corruption?
Partially. A stack overflow that overwrites the task control block or corrupts the scheduler will often cause a hang that the watchdog will recover via reset. However, the watchdog does not detect the overflow itself. For reliable detection, enable hardware stack checking, MPU guards, and FreeRTOS configCHECK_FOR_STACK_OVERFLOW. I treat the watchdog as the last line of defense; the first line is static allocation and stack high-water monitoring. If your supervisor is resetting frequently, check task stack margins before assuming a logic bug.
How do I handle firmware updates without triggering the watchdog?
Do not disable the watchdog during updates. Instead, either keep kicking it from your bootloader or update logic while proving progress, or reconfigure it to a longer timeout for the duration of the update. On Zephyr, the bootloader (MCUboot) can service the watchdog during flash writes. In FreeRTOS OTA, I break the flash erase/write into page-sized chunks and check in with the supervisor between chunks. If the update itself hangs, you still want the watchdog to reset and retry. A watchdog that is disabled for a 60-second update turns a recoverable glitch into a bricked device.