Embedded Debugging with JTAG and SWD: Tools, Techniques and Workflows

Debugging embedded targets through blinking LEDs and UART print statements will only get you so far. The moment you need to chase a hard fault that occurs once every few thousand wake cycles, or figure out why DMA is corrupting a buffer under load, you need direct access to the core. That means JTAG or SWD. In my experience, developers who invest early in a solid debug probe setup and a repeatable OpenOCD workflow save weeks of guesswork later. This article covers how both interfaces actually work, which tools are worth your money, how to wire and configure them reliably, and the day-to-day workflows I use to bring up new boards and trace elusive bugs on Arm Cortex-M and similar architectures.

JTAG Chain Mechanics vs SWD's Two-Wire Reality on Cortex-M

JTAG was defined by IEEE 1149.1 as a boundary-scan and test access method, long before it became the de facto embedded debug port. It uses four mandatory signals — TCK, TMS, TDI, TDO — plus an optional TRST. Every device on the chain has a TAP (Test Access Port) controller, and those TAPs are daisy-chained: TDI of one device feeds TDO of the next. That chaining is powerful for board-level testing with multiple FPGAs or MCUs, but it is also its biggest source of complexity.

SWD (Serial Wire Debug), introduced by Arm for the Cortex family, collapses that interface to just two signals: SWCLK and SWDIO, plus ground and optionally SWO for trace. It uses a packet-based protocol over a single bi-directional data line, with the host driving the clock. Inside the chip, SWD talks to a DAP (Debug Access Port) that bridges to the same AHB-AP and debug logic that JTAG would. For single-core Cortex-M0/M3/M4/M33/M7 parts, which dominate most IoT sensor nodes and wearables, SWD is all you need.

When JTAG Is Still Unavoidable

I've found that you still need full JTAG in three cases: multi-core SoCs where each core has its own TAP, legacy cores like Arm7TDMI or MIPS that predate SWD, and when you require boundary-scan for production testing. Some high-pin-count MCUs like the STM32H7 or NXP i.MX RT series expose both interfaces on the same pins and choose the protocol at power-on via a sequence on TMS/SWDIO. If you hold TMS high and toggle TCK, you stay in JTAG; the SWD switching sequence (at least 50 clocks with SWDIO high, then 0xE79E) moves the target into SWD mode.

Performance and Pin-Cost Trade-offs

From a pure bandwidth standpoint, 4-wire JTAG can clock faster in theory, but in practice SWD is more efficient because it has less turnaround overhead and most modern probes drive SWD at 10-25 MHz without issues. The real win is pin cost. On a 32-pin QFN where every GPIO matters, reclaiming two pins by using SWD instead of JTAG often decides the package choice. That is why almost every Cortex-M development board — ST Nucleo, Nordic nRF52 DK, Raspberry Pi RP2040 — routes only SWD.

Probe Selection: From ST-Link Clones to J-Link and CMSIS-DAP Realities

Your probe is the translator between USB and the target's debug port, and not all translators are equal. I've used everything from $5 ST-Link clones to $600 J-Link Pros, and the differences show up in stability, speed, and software support rather than raw ability to halt a core.

CMSIS-DAP is the open standard created by Arm. Any probe that implements it — from the on-board debugger on a BBC micro:bit to the cheap DAPLink adapters — will work with OpenOCD, pyOCD, and Keil. It is vendor-neutral and firmware-upgradeable, but historically limited to about 1-2 MB/s flash programming and no high-speed trace.

ST-Link/V2 and ST-Link/V3 are STMicroelectronics' probes. An ST-Link/V3SET can drive SWD at 24 MHz and includes a virtual COM port and power measurement, but it is optimized for STM32 parts. J-Link from SEGGER is the industry workhorse. Even the base J-Link EDU Mini handles almost every Cortex-M and RISC-V part, programs flash at 1-4 MB/s, and works flawlessly with GDB, OpenOCD, and Ozone. Where it really earns its price is in its GDB server stability and support for streaming trace.

Probe Interface Support Max SWD Clock OpenOCD Support Best For
ST-Link/V2 Clone SWD only 4 MHz Good (stlink driver) Low-cost STM32 prototyping
CMSIS-DAP (DAPLink) SWD, JTAG 10 MHz Excellent (cmsis-dap driver) Vendor-neutral, open-source boards
ST-Link/V3SET SWD, JTAG, SWO 24 MHz Good STM32H7/U5 high-speed debug
SEGGER J-Link BASE / EDU SWD, JTAG, SWO, ETM 15-25 MHz Excellent (jlink driver) Professional, multi-target work
J-Link PRO / Trace SWD, JTAG, 4-bit ETM Trace 25 MHz+ Limited (prefers J-Link GDB Server) Instruction trace on Cortex-M4/M7

In my lab, I keep a J-Link BASE for daily work and a handful of CMSIS-DAP adapters for field kits. If you are building a product around nRF or RP2040, pyOCD with CMSIS-DAP is perfectly comfortable. If you support multiple vendors, investing in one J-Link saves you from maintaining five vendor-specific drivers. One caution: avoid the cheapest ST-Link clones for new designs — many have outdated firmware that fails the SWD line reset sequence and will randomly drop the connection mid-flash.

Wiring and Signal Integrity Pitfalls That Kill Debug Sessions

More debug sessions fail due to wiring than toolchain bugs. SWD looks simple on a schematic, but the signals are fast enough that sloppy layout causes intermittent failures that look like software problems.

The Minimal Reliable Connection

For SWD you need: SWDIO, SWCLK, GND, NRST (optional but highly recommended), and 3.3V target sense (VTref). Do not power the target from the probe unless the probe explicitly supports it and you understand the current limit. Always connect ground — I use a 10-pin Cortex Debug Connector (0.05" pitch) even on tiny boards because it grounds multiple pins and locks mechanically. If you use Tag-Connect or pogo pins, keep the cable under 15 cm.

What I Check When OpenOCD Says "No Device Found"

First, measure VTref at the probe header. If it is floating or 1.8V when your core is 3.3V, the probe's level shifters are misconfigured. Second, check pull resistors. SWDIO needs an internal pull-up (most MCUs provide this), SWCLK needs a pull-down. I've seen boards where an external 10k pull-up on SWCLK held the clock high enough that the probe could not drive a clean low — removing it fixed the issue. Third, look at NRST. Some probes drive NRST open-drain; if your reset line has a strong pull-up and a capacitor larger than 100 nF, the probe cannot assert reset fast enough. Add a 100 ohm series resistor to SWCLK/SWDIO near the MCU and keep trace length matched and short.

For JTAG chains, verify TDO->TDI ordering and that TRST is either pulled high or driven. A floating TRST will leave a TAP in reset and break the entire chain. On mixed-voltage boards, use a proper level translator like the LSF0108 rather than resistor dividers — dividers round off clock edges and cause setup violations above 4 MHz.

OpenOCD Configuration That Actually Connects on the First Try

OpenOCD (Open On-Chip Debugger) is the glue between your probe and GDB. Once you understand its two-part configuration — interface and target — most connection issues disappear. In my experience, the fastest way to debug OpenOCD is to run it with -d3 and watch the SWD acknowledgment packets.

A minimal working setup for an STM32G4 with an ST-Link looks like this:

# interface configuration
source [find interface/stlink.cfg]
transport select hla_swd

# target configuration
source [find target/stm32g4x.cfg]

# Optional: increase adapter speed after init
adapter speed 4000

init
reset halt

For a CMSIS-DAP probe on an nRF52840, swap the interface:

source [find interface/cmsis-dap.cfg]
transport select swd
adapter speed 10000
source [find target/nrf52.cfg]
init
reset halt

The order matters: interface first, then transport, then target. The hla_swd transport is specific to ST-Link's high-level adapter; plain swd is used for CMSIS-DAP and J-Link. If OpenOCD reports Error: open failed, it is a USB permission or driver issue — on Linux, install the 99-openocd.rules udev file. If it reports Error: target not examined yet, your reset_config is wrong. For boards that lack a dedicated reset line, add reset_config srst_none or reset_config connect_assert_srst for devices that need reset asserted during connect, like some NXP parts.

When bringing up a brand new chip, I start with the lowest adapter speed (100 kHz) and verify the IDCODE. For SWD, OpenOCD will print the DPIDR (Debug Port ID Register). Compare it to the reference manual — for example, Cortex-M4 DPIDR is typically 0x2BA01477. If you read 0x00000000 or 0xFFFFFFFF, your wiring or power is wrong, not your config. The Zephyr Project Documentation has excellent examples of west-driven OpenOCD invocations for dozens of boards that you can copy as a starting point.

Flashing and Verifying Without an IDE

Once connected, you can flash directly from the OpenOCD telnet interface on port 4444:

halt
flash write_image erase build/firmware.elf
verify_image build/firmware.elf
reset run
exit

I prefer ELF over binary because it carries the load address. For production programming, add flash bank definitions manually if your target config does not set them. This also lets you lock flash after programming, which ties into securing the debug port later.

Live Debugging Workflows: Breakpoints, Watchpoints, and ITM Tracing

Hitting a breakpoint is easy; using breakpoints without disturbing timing is harder. Cortex-M cores provide 6-8 hardware breakpoints (via the FPB unit) and 4 watchpoints (via DWT). Software breakpoints work by patching flash with a BKPT instruction, but they require flash reprogramming and stall. Hardware breakpoints work by comparing the program counter without modifying memory and are essential for debugging code in flash or time-critical ISRs.

My typical GDB session after launching OpenOCD's GDB server (port 3333) is:

arm-none-eabi-gdb build/firmware.elf
(gdb) target extended-remote :3333
(gdb) monitor reset halt
(gdb) load
(gdb) monitor arm semihosting enable
(gdb) break main
(gdb) continue

From there, use watch myBuffer[32] to set a watchpoint that halts only when that memory location is written — invaluable when tracking down stack overflows or DMA overruns. The Memory Management in Embedded C: Static Allocation, Pool and Arena Patterns discussion of pool allocators is directly relevant here, because watchpoints let you catch the exact instruction that writes past a pool block boundary.

Real-Time Insight Without Halting

Halting the core destroys real-time behavior. For interrupt-heavy firmware, I rely on non-intrusive trace. ITM (Instrumentation Trace Macrocell) sends printf-style data over the single SWO pin without halting. Configure it once and you get timestamped logs at near-zero overhead.

On an STM32, enabling ITM in firmware looks like enabling the trace clock and stimulus port 0. Then in OpenOCD:

tpiu config internal itm.fifo uart off 168000000
itm ports on 0x1

This streams ITM port 0 to a file that you can tail with tail -f itm.fifo. SEGGER's RTT (Real Time Transfer) is even faster because it uses a RAM buffer and the debug access port instead of SWO, achieving hundreds of KB/s without extra pins. If you are debugging Interrupt Handling Best Practices: Priority, Latency and ISR Design issues where halting would mask a race condition, RTT or ITM is the only way to see what is happening without perturbing timing.

RTOS-Aware Debugging

When an RTOS Fundamentals: FreeRTOS vs Zephyr for Embedded Projects stack is running, raw thread stacks are meaningless. Load the RTOS plugins for GDB: OpenOCD ships with FreeRTOS-openocd.c helpers that parse the TCB lists. Add source [find rtos/FreeRTOS.tcl] to your OpenOCD target config and GDB will show thread names, priorities, and which task is blocked on a queue. For Zephyr, the west debug wrapper automatically loads the correct Python pretty-printers. The FreeRTOS Documentation covers the configUSE_TRACE_FACILITY settings you must enable to expose those kernel structures.

Recovering Bricked Targets and Securing the Debug Port in Production

Every embedded engineer eventually bricks a board — usually by configuring SWD pins as GPIO, entering deep sleep with debug disabled, or enabling readout protection (RDP) incorrectly. Recovery requires understanding how the core boots.

If the firmware reconfigures PA13/PA14 (the SWD pins on STM32) as outputs, the debugger will not be able to connect after reset release because the pins are no longer in debug mode for more than a few milliseconds. The fix is to assert reset, connect under reset, and halt before the firmware reconfigures the pins. In OpenOCD this is reset_config srst_only connect_assert_srst followed by init; reset halt. Some probes need you to hold the physical reset button while issuing openocd -f interface/... -f target/... and release it a half-second later.

Low-power modes are similar. On STM32L4, setting DBGMCU_CR_DBG_SLEEP is required or the core powers down the debug logic in Stop mode. If you forgot it, you can still recover by connecting while holding reset and mass-erasing. The command stm32l4x mass_erase 0 or nrf52_recover for Nordic parts will wipe flash and remove protection, at the cost of your firmware.

Locking the Door Before Shipping

An open SWD port in the field is a security hole — anyone can dump firmware and extract keys. All modern MCUs provide some form of access control. STM32 offers RDP levels: Level 0 is open, Level 1 prevents flash readout via debugger but allows reprogramming, Level 2 permanently disables debug (irreversible). nRF52 uses APPROTECT and UICR. I recommend Level 1 for development units that may need field updates, and Level 2 only after thorough testing where you have a UART or USB DFU fallback. Always test the locking sequence on a disposable board first — I've seen teams permanently lock engineering samples because they set RDP Level 2 before verifying their bootloader.

For production programming stations, use a scripted OpenOCD flow that flashes, verifies, sets option bytes, and then immediately disconnects. Never leave SWO or RTT logging enabled in release builds; those channels can leak sensitive data even if flash readout is disabled.

Frequently Asked Questions

Can I use SWD and JTAG on the same header without rewiring?

Yes, if your MCU supports both on shared pins. Most Cortex-M parts multiplex JTAG and SWD on the same five pins (TCK/SWCLK, TMS/SWDIO, TDI, TDO, TRST). You can wire a standard 10-pin or 20-pin Cortex connector that carries all signals and let the probe select the protocol via software. Just ensure your firmware does not reconfigure TDI/TDO as GPIOs if you need JTAG later, and add the correct pull resistors so the switching sequence is reliable.

Why does OpenOCD connect once and then fail with "SWD DP error" after flashing?

That usually means your firmware immediately enters a low-power mode, disables the debug power domain, or reconfigures SWD pins. The target is alive but the DAP is powered down, so the probe gets a FAULT response. Enable debug in sleep/stop in your startup code (for STM32 set DBGMCU_CR, for other vendors see their debug retention registers) or add a few hundred milliseconds delay at boot before entering sleep so the debugger can halt the core first.

Is SWO/ITM trace available on all Cortex-M chips?

No. Cortex-M0 and M0+ have very limited or no ITM/ETM — they expose only basic SWD. Cortex-M3, M4, M7, and M33 include ITM and DWT, and some include ETM for full instruction trace. You also need a probe that supports SWO capture (ST-Link/V3, J-Link, CMSIS-DAP v2). Check your reference manual for the TPIU and ITM chapters; if TPIU is absent, fall back to RTT which only needs RAM and SWD.

How many hardware breakpoints can I actually use with GDB?

Six on most Cortex-M3/M4 (FPB v1) and eight on Cortex-M7/M33 (FPB v2), but GDB reserves one for its own use. If you set more than the available number, GDB will silently fall back to software breakpoints which require reprogramming flash and may fail on write-protected sectors. Use the command info breakpoints to see the type (hw vs sw) and remove unused ones. Watchpoints are separate — you get four comparators via DWT, each can monitor up to 4 bytes.

Related Articles

References & Standards: FreeRTOS Documentation · Zephyr Project Documentation · MQTT Specification