Real-time bugs don't crash your system politely. They corrupt a queue at 3 AM after 72 hours of uptime, they shift a PWM edge by 4 microseconds when the Wi-Fi radio transmits, they deadlock two high-priority tasks for exactly one scheduler tick. In my experience, the moment you try to catch these faults with a standard breakpoint, you destroy the timing that created them. Halting the core hides the race condition, masks the jitter, and leaves you staring at a system that behaves perfectly while under observation. That's why serious embedded work demands a different toolbox: one that observes without intruding.
This article covers the three pillars I rely on when debugging live RTOS-based firmware - JTAG for invasive but controlled inspection, hardware trace for nanosecond-accurate execution history, and logic analyzers for correlating the outside world with what the CPU thinks it's doing. We'll look at how these tools work together on modern ARM Cortex-M parts (STM32H7, nRF53, NXP i.MX RT) and how to configure them to catch the faults that printf debugging never will.
Why Real-Time Bugs Evade Conventional Breakpoint Debugging
On a bare-metal super-loop, a breakpoint is usually harmless. On an RTOS, it's destructive. When you hit a breakpoint on Cortex-M, the core halts, but peripherals keep running. SysTick continues to count, DMA transfers complete, watchdogs expire, and communication peripherals overflow. By the time you hit "continue," you've introduced a timing anomaly that can be worse than the original bug. I've found that this is the single biggest reason teams waste weeks chasing phantom issues that only appear in the field.
The classic example is a priority inversion that only manifests under load. Task A (high priority) waits on a mutex held by Task C (low priority), while Task B (medium priority) starves Task C. If you halt Task A to inspect its stack, Task B stops being scheduled, Task C suddenly gets CPU time, releases the mutex, and the problem vanishes. You resume and conclude there is no bug.
The Heisenberg Effect in RTOS Context Switching
Every halt-based debug action perturbs the scheduler. The FreeRTOS Documentation is explicit that kernel-aware debugging requires special handling because the scheduler's ready lists and TCBs are in flux during a context switch. A debugger that reads memory while PendSV is pending can show you a corrupted list that was never actually corrupted - you just sampled it mid-update.
This is why I treat breakpoints as a tool for bring-up and initialization bugs only. For anything involving concurrency, timing, or interrupts, I switch to non-intrusive observation. That means instruction trace, data watchpoints, and external capture. The goal is to record what happened and analyze it after the fact, not to freeze time and hope the bug stays still.
Classifying Faults That Demand Trace
In my lab, I categorize trace-dependent bugs into three buckets. First, temporal violations: missed deadlines, jitter on control loops, ISR latency spikes. Second, data-dependent races: corrupted shared buffers, use-after-free across tasks, and stack overflows that only overflow under a specific interrupt nesting. Third, interaction faults: where the firmware is logically correct but the interaction with external hardware - a sensor's busy line, an SPI flash's timing, a radio's coexistence signal - violates assumptions.
If your bug falls into any of these, a logic analyzer and a trace probe will save you days. If you rely solely on halt-mode JTAG, you'll be guessing.
JTAG Under the Hood: From TAP Controller to Live Memory Inspection
JTAG is often dismissed as just "the thing you use to flash and set breakpoints," which dramatically undersells it. The IEEE 1149.1 Test Access Port (TAP) gives you a state machine-driven backdoor into the core. On ARM Cortex-M, JTAG (or Serial Wire Debug, SWD) connects to the Debug Access Port (DAP), which in turn gives access to the AHB-AP for memory access without necessarily halting the core.
Understanding this matters because it unlocks live memory inspection. With a decent probe (Segger J-Link, ST-LINK v3, or even a CMSIS-DAP build with OpenOCD), you can read RTOS structures, peripheral registers, and RAM buffers while the target runs at full speed. The trick is using the AHB-AP correctly and not accidentally triggering a halt.
Configuring OpenOCD for Non-Halting Access
Most engineers use their IDE's default OpenOCD or pyOCD script and never touch the configuration. I've found that tuning it is essential for stable real-time work. You want to enable memory access while running and ensure your connect sequence doesn't reset the target when you attach.
# stm32h7-rtos-attach.cfg - attach without reset, allow live memory reads
source [find interface/stlink.cfg]
transport select hla_swd
source [find target/stm32h7x.cfg]
# Do not assert SRST on connect - critical for field-debugging
reset_config none
init
halt # optional - remove if you want to attach to running target
# Enable AHB-AP access while running
arm dap apreg 1 0x04 0x00000000
# Set CSW to enable master access while core runs
Once attached, you can script GDB to dump task states periodically. I keep a small Python GDB script that reads the pxCurrentTCB, pxReadyTasksLists, and xDelayedTaskList1 from FreeRTOS every 100ms without halting. That alone has shown me starvation patterns that breakpoints hid.
Data Watchpoints via the DWT Unit
The Data Watchpoint and Trace (DWT) unit is part of the Cortex-M debug architecture and accessible via JTAG/SWD. Unlike a breakpoint that halts on code execution, a watchpoint can halt (or more usefully, trigger trace) on data access. For debugging heap corruption or a buffer overflow, I configure a watchpoint on the last 4 bytes of a task stack or a queue structure.
// Configure DWT comparator 0 to watch for writes to a suspect queue
// Address: &xQueue->pcWriteTo, Mask: 0 (match exact address), Function: watchpoint on write
#include "core_cm7.h"
void dwt_watch_queue_corruption(QueueHandle_t q) {
uint32_t *watch_addr = (uint32_t*)&((Queue_t*)q)->pcWriteTo;
// Disable comparator while configuring
DWT->COMP0 = (uint32_t)watch_addr;
DWT->MASK0 = 0; // No address masking
// FUNCTION = 0x05: Data write watchpoint, emit event without halting if needed
// For halting watchpoint: 0x06
DWT->FUNCTION0 = (0x05 << 0) | (1 << 11); // DATAVSIZE = word
// Enable DWT and trace
CoreDebug->DEMCR |= CoreDebug_DEMCR_TRCENA_Msk;
DWT->CTRL |= DWT_CTRL_CYCCNTENA_Msk;
// Route to TPIU/ITM for non-halting trigger, or halt via DEMCR
// For field use, I prefer non-halting + ITM event
}
This catches the exact instruction that smashes your queue, with the full register context, without needing to guess where to put a breakpoint. Pair it with the Fault Status Registers (CFSR, HFSR, MMAR) and you have a precise post-mortem path for hard faults.
Instruction Trace Decoded: ETM, ITM and SWO in Practice
If JTAG lets you inspect state, trace lets you replay time. On Cortex-M3/M4/M7/M33/M55, you get several trace sources: Embedded Trace Macrocell (ETM) for full instruction trace, Instrumentation Trace Macrocell (ITM) for software events, and Data Watchpoint and Trace (DWT) for periodic counters and watchpoint events. All of this funnels through the Trace Port Interface Unit (TPIU) and out via either a 4-bit parallel trace port or the single-pin Serial Wire Output (SWO).
In my experience, SWO is underutilized. It only needs one pin (SWO/TDO) and a ground, runs at 1-8 MHz, and gives you ITM printf-style logging at microsecond resolution with almost zero CPU overhead - a single store to the ITM stimulus port. ETM is more powerful but needs a high-end trace probe (Segger J-Trace, Lauterbach) and 4-5 dedicated pins, which many IoT boards don't route.
Getting ITM and SWO Running Without an IDE
Most vendor IDEs hide the setup, but when you're on custom hardware, you need to configure it yourself. The sequence is: enable trace in CoreDebug, configure TPIU/SWO baud rate, unlock and enable ITM, and enable stimulus ports. This works whether you are on Zephyr, FreeRTOS, or bare-metal.
// Minimal ITM/SWO init for STM32H7 at 400MHz core, 2MHz SWO
// Adapt TPIU_ACPR for your core clock / SWO baud
void itm_swo_init(void) {
// Enable trace subsystem
CoreDebug->DEMCR |= CoreDebug_DEMCR_TRCENA_Msk;
// TPIU: Select SWO protocol, set prescaler
// ACPR = (core_clock / swo_baud) - 1
TPIU->ACPR = (400000000 / 2000000) - 1;
TPIU->SPPR = 2; // 0x02 = SWO Manchester/NRZ
TPIU->FFCR = 0x102; // Enable formatter
// Lock access for ITM
ITM->LAR = 0xC5ACCE55;
ITM->TCR = (1 << 0) | // ITMENA
(1 << 1) | // TSENA (timestamp)
(1 << 3) | // SYNCENA
(3 << 10); // TSPrescale
ITM->TER = 0x1; // Enable stimulus port 0
// Also enable DWT cycle counter for timestamps
DWT->CTRL |= DWT_CTRL_CYCCNTENA_Msk;
}
// Ultra-low overhead trace - 1-2 cycles if ITM FIFO not full
static inline void trace_task_switch(uint8_t from_id, uint8_t to_id) {
if (ITM->PORT[0].u32 != 0) return; // Check FIFO ready
ITM->PORT[0].u32 = (0xA5 << 24) | (from_id << 8) | to_id;
}
With that running, you can stream task switch events, ISR entry/exit, and custom markers directly from the RTOS context switch hook. On FreeRTOS, I hook traceTASK_SWITCHED_IN() and traceTASK_SWITCHED_OUT() to emit a 4-byte packet. On Zephyr, I use the tracing subsystem via Zephyr Project Documentation tracing hooks. The overhead is under 0.5% CPU at 10k switches per second - far less than UART logging, which can easily add 5-10% and perturb timing.
When to Step Up to ETM
ITM tells you what you instrumented. ETM tells you everything. It records every branch, exception, and cycle count, allowing you to reconstruct exact execution flow without any code changes. I've used ETM to find a subtle bug where a compiler optimization reordered a volatile access inside a critical section. ITM wouldn't have shown it because I didn't think to instrument that line.
The downside is bandwidth and tooling. A Cortex-M7 at 480 MHz can generate 200-400 MB/s of ETM data during dense code. You need a 4-bit ETR (Embedded Trace Router) or external trace buffer, and a debugger that can decode it. For most IoT products, I reserve ETM for lab boards where I've routed the trace pins. In production hardware, SWO + DWT is the pragmatic choice. For more on keeping timing deterministic while tracing, see the companion discussion on FreeRTOS Task Scheduling: Priorities, Preemption and Time Slicing.
Correlating Logic Analyzer Captures with RTOS Task Execution
Hardware trace shows you what the CPU did. A logic analyzer shows you what the rest of the world did. The real insight comes from correlating the two. I never debug a timing-sensitive driver - SPI, I2C, I2S, or 802.15.4 radio coexistence - without both captures time-aligned.
The technique is simple: dedicate 2-3 GPIO pins as a trace port. Toggle them at key RTOS events, and capture those same pins on the logic analyzer alongside your external signals. The GPIOs become a low-jitter, hardware-timed marker stream that ties software state to electrical reality.
Wiring GPIOs as a Software Logic Analyzer Channel
Choose GPIOs on the same port for single-cycle writes via BSRR. I typically use a 3-bit code: one bit for ISR context, two bits for task ID. That gives 8 states, enough to distinguish idle, high-priority control loop, and communication tasks. The write must be atomic and fast - no HAL overhead.
On a recent motor controller project, we had a jitter issue where the FOC loop occasionally ran 12 microseconds late. ITM showed the delay, but not the cause. By routing the FOC task marker to Channel 7 of a Saleae Logic Pro 16 and capturing the quadrature encoder index pulse on Channel 0, we saw the encoder ISR preempted the FOC task, but only when the BLE radio's coexistence line (Channel 1) was high. The RTOS tracing alone missed the radio interaction entirely because it was an external signal. Without the analyzer, we'd still be tuning scheduler priorities.
Choosing Depth, Sampling Rate and Trigger
For RTOS correlation, I run the analyzer at 25-50 MS/s even if my signals are only 1 MHz. The extra resolution lets you measure ISR latency to 40ns. Depth matters more than speed; with 500M samples you can capture several seconds around a rare fault. Set the trigger on the GPIO marker that represents the fault condition - for example, trigger when Task C's marker stays high longer than 2ms, indicating a deadline miss. Modern analyzers like Sigrok-compatible devices or Saleae allow protocol analyzers to decode I2C/SPI while still showing the task markers in parallel.
One practical tip: keep the GPIO marker writes in RAM-resident functions (place them in ITCM or use __attribute__((section(".ramfunc")))) so flash wait states don't add jitter to your markers. I've measured 6-cycle variance just from flash cache misses on marker pins.
Non-Intrusive Instrumentation: RTT, DWT and Post-Mortem Trace Buffers
When you can't attach a probe in the field, you need on-device traces. Two techniques have proven invaluable for me: SEGGER RTT (Real Time Transfer) for high-bandwidth, non-blocking logging, and in-RAM post-mortem buffers that survive resets.
RTT uses a small ring buffer in RAM that the debugger polls via the AHB-AP while the core runs. Unlike semihosting or UART, it doesn't block, doesn't need interrupts, and can push 1-2 MB/s with ~1 microsecond per write. I instrument error paths and timing-critical sections with RTT and leave it enabled in beta builds. The key is to make the buffer large enough (4-8KB per channel) and to handle overflow gracefully - drop oldest, not newest.
Cycle-Accurate Profiling with DWT_CYCCNT
The DWT cycle counter is a free, 32-bit, 1-cycle resolution timestamp. For profiling ISR latency or task execution time, it's unbeatable. I wrap critical sections to measure actual execution time and detect violations without stopping the core.
This pairs well with memory protection. If you are seeing intermittent heap corruption, combining DWT watchpoints with the approach in RTOS Memory Management: Static Allocation, Pools and Heap Strategies often reveals a wild pointer from a fragmented heap. Static pools reduce the search space dramatically.
RAM Trace Buffers That Survive Watchdog Resets
For field returns, you won't have a trace probe. I reserve 16-32KB of no-init RAM (place it in a section excluded from startup clearing) and write a circular buffer of compact events: timestamp (DWT_CYCCNT), task ID, event code, and one data word. On a hard fault or watchdog reset, the buffer persists. The reset handler dumps it to flash or RTT on next boot.
The structure is deliberately small - 8 bytes per event - so you can store ~4000 events in 32KB. That's enough for several seconds of history around a crash. I've solved multiple brown-out and EMI-induced crashes this way where the device reset before any debugger could attach. If you also use the watchdog as your last line of defense, align your buffer dump with the strategy in Watchdog Timers: Building Reliable Embedded Systems That Self-Recover - dump before petting, or before the reset, not after.
Building a Deterministic Debugging Workflow for Hard Faults and Race Conditions
Tools alone don't solve bugs; workflow does. My deterministic workflow has four stages, and I follow them in order to avoid wasting time on red herrings.
First, reproduce with observation only. Don't fix, don't add delays. Attach SWO/RTT and capture a logic analyzer trace around the failure. If you can't reproduce on the bench, ship an instrumented beta with the RAM trace buffer enabled.
Second, narrow with watchpoints and triggers. Once you have a timestamp for the fault from ITM or the analyzer, set a DWT watchpoint or an analyzer trigger to stop just before it. For hard faults, read CFSR, BFAR, and the stacked PC/LR first - they tell you if it was a bus fault, precise or imprecise, and which instruction faulted.
Third, isolate scheduling. Use trace to confirm task states at the moment of failure. Is the faulting task actually the one with the highest priority? Is an ISR holding a resource too long? I've seen SPI DMA completion ISRs that copy 256 bytes in the ISR, adding 40 microseconds of jitter to the entire system. Moving that copy to a high-priority task cut jitter by 80% and disappeared from trace.
Finally, fix and verify with trace, not just tests. After a fix, re-run the same trace capture and compare cycle counts, jitter histograms, and execution ordering. A fix that passes functional tests but adds 3% CPU overhead or increases worst-case latency by 5 microseconds is not a fix for a real-time system. Trace gives you that quantitative proof.
The investment pays off. On a recent Zephyr-based sensor hub, we reduced rare field crashes from 1 in 10k hours to none in 100k hours after adding ITM markers to all IPC paths and a logic analyzer capture of the sensor's data-ready and power-enable lines. The root cause was a 2ms window where the sensor could NACK while the driver assumed ACK - invisible in code review, obvious when the trace and analyzer were overlaid. The RTOS Synchronization: Mutexes, Semaphores and Event Groups primitives were used correctly; the bug was in the assumption about external timing, which only showed up when software and electrical views were merged.
| Technique | Intrusion / Overhead | Best For | Hardware Requirements |
|---|---|---|---|
| Halt-Mode JTAG/SWD Breakpoints | High - Halts core, disturbs RTOS timing | Bring-up, init code, single-threaded bugs | Any SWD probe, 2 pins |
| ITM / SWO + DWT | Very Low - 1-2 cycles per event | Task switches, ISR latency, lightweight logging | SWD + SWO pin, SWO-capable probe |
| ETM Instruction Trace | Zero - Full hardware capture | Compiler bugs, precise execution reconstruction | 4-bit trace port + high-end probe (J-Trace) |
| RTT + RAM Trace Buffer | Low - ~1 µs per packet, no halt | Field diagnostics, post-mortem crashes | No extra pins, RAM reservation only |
| Logic Analyzer + GPIO Markers | Minimal - Single BSRR write (~4 cycles) | Hardware-software correlation, jitter analysis | 2-3 GPIOs + external analyzer at 25+ MS/s |
Frequently Asked Questions
Can I use SWO trace on a low-pin-count IoT board with no dedicated trace header?
Yes, if you can spare the SWO pin. On most Cortex-M designs SWO is multiplexed with TDO/SWO on the SWD header. You don't need a parallel trace connector. Enable SWO at 1-2 MHz, connect your probe's SWO line (often the same as TDO on 10-pin Cortex Debug connectors), and you get ITM and DWT trace over one wire. I've run SWO successfully on 48-pin STM32G4 and nRF52840 boards with just the standard 2x5 SWD header and a Segger J-Link EDU. Keep the wire short - under 10cm - for stable signaling at higher baud rates.
Why does my logic analyzer show GPIO markers with jitter even though my task is high priority?
Three common causes: flash wait states, interrupt preemption, and bus arbitration. If your marker function lives in flash, instruction fetch stalls add 2-8 cycles of variance. Move it to RAM/ITCM. Second, a higher-priority ISR can preempt between your marker write and the actual event - measure ISR nesting with ITM to confirm. Third, on Cortex-M7 with AXI bus, a concurrent DMA transfer can arbitrate the GPIO peripheral bus. Use the DWT CYCCNT timestamp alongside the GPIO to isolate which source dominates. In my experience, moving markers to RAM cuts jitter from ~120ns to ~20ns on an H7.
How do I debug a hard fault that only happens after days of uptime?
Don't try to catch it halting. Use a persistent RAM trace buffer plus fault register capture. Reserve a no-init RAM section that the startup code doesn't clear, continuously log compact events (timestamp, task ID, fault status), and on HardFault, copy CFSR, HFSR, BFAR, MMFAR, and the stacked registers to that buffer before resetting. Also configure the watchdog to trigger a reset with a dump. On next boot, write the buffer to flash or stream over RTT. This gives you the exact instruction and task state days later without needing a debugger attached. Pair it with DWT watchpoints on suspect buffers to narrow the window.
Is JTAG/SWD debugging safe to leave enabled in production firmware?
Functionally safe but security-sensitive. Leaving SWD enabled exposes memory read/write via the DAP. For development and beta, I leave it enabled with a strong option-byte setting that can be locked later. For production, either disable SWD pins via GPIO reconfiguration after boot, or use the MCU's readout protection (RDP level 1/2 on STM32, APPROTECT on nRF53). Note that RDP will also block legitimate field debugging, so consider a challenge-response scheme or leave a secure bootloader path to temporarily re-enable debug. Never ship with RDP level 0 if the device handles sensitive data.