Interrupt handling is one of those areas where embedded theory and board-level reality collide. On paper, the processor refines an interrupt in a few cycles, pushes context, and vectors to your handler. In practice, I've seen systems miss critical sensor deadlines because a UART ISR held the CPU for 400 microseconds servicing a FIFO, or because two interrupts were assigned the same priority when one needed strict preemption. Getting interrupts right is less about writing handler functions and more about architecting how your system responds to asynchronous events under worst-case load. Over the last decade working on Cortex-M based motor controllers, industrial gateways, and low-power sensor nodes, I've learned that the most reliable systems treat interrupt latency as a budgeted resource, keep interrupt service routines (ISRs) brutally short, and defer everything else to deterministic background processing.
Why Your Interrupt Latency Budget Determines System Determinism
Interrupt latency — the time from the assertion of an interrupt signal to the execution of the first instruction of your ISR — is the foundation of real-time responsiveness. It is not a single number from a datasheet. On an ARM Cortex-M4, the core advertises 12 cycles of stacking and vector fetch, but total latency includes synchronizer delays, bus arbitration, Flash wait states if your vector table lives in Flash, and the time spent completing a current instruction or a higher-priority ISR. In my experience, the number that matters is worst-case latency under full system stress, not typical latency measured on an idle bench.
I start every new project by defining a latency budget for each time-critical interrupt. For example, on a BLDC motor controller, the current-loop ADC must be serviced within 2µs of the PWM sync pulse to keep the field-oriented control stable. That gives me a hard deadline. I then work backward: total budget = hardware synchronizer (1-2 cycles) + instruction completion (up to 10 cycles on Cortex-M with load/store) + stacking (12 cycles) + Flash prefetch stall (0-6 cycles) + longest critical section with interrupts disabled. If that sum exceeds 2µs at 168MHz (~336 cycles), I have no margin. The only levers are shortening critical sections, moving the vector table to SRAM or ITCM, and ruthlessly prioritizing that interrupt above all others.
Calculating Worst-Case Latency With Critical Sections
The biggest hidden contributor is code that disables interrupts. Every __disable_irq() or entry into a FreeRTOS critical section (taskENTER_CRITICAL()) adds directly to the worst-case latency of all interrupts at or below the mask level. The FreeRTOS Documentation is explicit about this: FreeRTOS masks interrupts up to configMAX_SYSCALL_INTERRUPT_PRIORITY using BASEPRI, not globally, specifically to bound latency for high-priority interrupts. I've found that choosing which interrupts are allowed to use FreeRTOS API calls versus which run above that threshold is the single most important RTOS decision. On one industrial I/O project, we kept the encoder index interrupt at priority 0 (never masked), the RTOS-aware comms interrupts at priority 5-6, and everything else lower. That guaranteed the encoder never jittered, even when the TCP stack held a critical section for 15µs.
A practical technique I use is to instrument DWT_CYCCNT or a free-running timer to measure disabled time. Toggle a GPIO at entry and exit of critical sections during testing and capture with a logic analyzer. You will often find one utility function that disables interrupts for far longer than you assumed, typically during a flash write or a poorly implemented ring buffer.
Decoding Nested Vectored Interrupt Controller Priority Grouping on ARM Cortex-M
The Cortex-M NVIC looks simple until you have to configure priority grouping. The 8-bit priority register is split into preemption priority and subpriority (sometimes called group vs. sub-group) via the PRIGROUP field in SCB->AIRCR. Preemption priority decides if an ISR can preempt another active ISR. Subpriority only decides the order pending interrupts of the same preemption level will be serviced — it does not cause preemption. Misunderstanding this distinction has caused more field bugs in my experience than any other NVIC configuration error.
On STM32, for instance, you have 4 bits of implemented priority (0-15). If you set NVIC_PRIORITYGROUP_2 (2 bits preemption, 2 bits subpriority), you get 4 preemption levels (0-3) and 4 subpriorities within each. Two interrupts with preemption 1, subpriority 0 and 1 will not preempt each other; the lower subpriority number simply goes first if both become pending while a higher preemption ISR is running. For most real-time systems, I've found that grouping with all bits as preemption priority (NVIC_PRIORITYGROUP_4 or NVIC_PRIORITYGROUP_0 depending on vendor naming) is the safest default. You want explicit preemption control, not tie-breaking.
Configuring Priorities Without Guessing
Never assign priorities sequentially. Group them by urgency class. In my current template, I use:
Priority 0-1: Hard real-time, never masked, never calls RTOS APIs (e.g., PWM fault, encoder capture, DMA error). Priority 2-4: Fast RTOS-aware interrupts that signal work (e.g., ADC conversion complete, high-rate sensor SPI DMA). Priority 5 and lower (numerically higher): Slow peripherals (UART, I2C, button GPIO). This aligns with the FreeRTOS requirement that any ISR calling FromISR APIs must be at or below configMAX_SYSCALL_INTERRUPT_PRIORITY. The Zephyr Project Documentation describes a similar model with its irq_lock() and priority ceiling for kernel-aware ISRs.
// STM32 HAL example: Consistent priority grouping and assignment
// Call this once at boot before any IRQ enable
HAL_NVIC_SetPriorityGrouping(NVIC_PRIORITYGROUP_4); // 4 bits preemption, 0 subpriority
// Hard real-time, not masked by RTOS. Never calls FreeRTOS API.
HAL_NVIC_SetPriority(ADC1_2_IRQn, 0, 0);
HAL_NVIC_SetPriority(TIM1_UP_TIM10_IRQn, 1, 0); // PWM sync
// RTOS-aware interrupts. Must be >= configLIBRARY_MAX_SYSCALL_INTERRUPT_PRIORITY (e.g., 5)
HAL_NVIC_SetPriority(DMA1_Channel1_IRQn, 5, 0);
HAL_NVIC_SetPriority(USART1_IRQn, 6, 0);
HAL_NVIC_SetPriority(EXTI9_5_IRQn, 7, 0);
// Verify: any priority numerically lower than 5 must not call xQueueSendFromISR
configASSERT((5 << (8 - __NVIC_PRIO_BITS)) == configMAX_SYSCALL_INTERRUPT_PRIORITY);
If you are debating whether to use an RTOS at all, this priority architecture is central to that choice. I often reference the trade-offs discussed in Bare Metal vs RTOS: When Each Approach Makes Sense — bare metal gives you absolute control over NVIC and global interrupt masks, while an RTOS imposes a ceiling (BASEPRI) that simplifies reasoning but requires disciplined assignment. And if you are comparing RTOS options, RTOS Fundamentals: FreeRTOS vs Zephyr for Embedded Projects provides a good breakdown of how each kernel handles interrupt masking internally.
Keeping ISRs Lean: The Flag, Queue, and Deferred Handler Pattern
The cardinal rule I teach junior engineers is: an ISR does the minimum work needed to acknowledge the hardware and record that the event happened. Nothing more. Polling a sensor, formatting a string, or walking a linked list inside an ISR is how you create jitter for every other interrupt in the system.
In my experience, the most robust pattern is a two-stage design: the ISR captures time-critical data and defers processing. The second stage can be a high-priority task, a work queue, or a bare-metal main-loop handler. The ISR's jobs are: 1) clear the interrupt source, 2) capture volatile data (e.g., timer capture register, ADC result) into a lock-free structure, 3) wake the deferred handler, and 4) exit.
What Belongs Inside the ISR Versus Deferred Work
Take a quadrature encoder timer overflow ISR. The only thing that must happen at interrupt time is latching the timer count and direction bit before it changes again. Calculating position delta, checking for faults, or updating PID belongs deferred. For a UART RX interrupt with a FIFO, the ISR should drain the FIFO into a ring buffer; parsing an AT command frame should happen later.
#define RING_SIZE 128
typedef struct {
volatile uint16_t head;
volatile uint16_t tail;
uint8_t buf[RING_SIZE];
} ring_buf_t;
static ring_buf_t uart_rx_ring;
static TaskHandle_t uart_task_handle;
// Minimal UART ISR - no parsing, no printf, no blocking
void USART1_IRQHandler(void)
{
BaseType_t xHigherPriorityTaskWoken = pdFALSE;
// Must read SR/DR to clear RXNE on STM32; check your reference manual
while (USART1->SR & USART_SR_RXNE) {
uint8_t data = (uint8_t)USART1->DR;
uint16_t next_head = (uart_rx_ring.head + 1) % RING_SIZE;
if (next_head != uart_rx_ring.tail) {
uart_rx_ring.buf[uart_rx_ring.head] = data;
uart_rx_ring.head = next_head;
} else {
// Ring full: count overrun, do NOT block
// Handle error counter in deferred context
}
}
// Wake deferred task directly - no queue overhead if only signaling
vTaskNotifyGiveFromISR(uart_task_handle, &xHigherPriorityTaskWoken);
// Request context switch if deferred task is higher priority than current
portYIELD_FROM_ISR(xHigherPriorityTaskWoken);
}
// Deferred handler runs as a high-priority task, with interrupts enabled
void vUARTTask(void *pvParameters)
{
for (;;) {
ulTaskNotifyTake(pdTRUE, portMAX_DELAY); // Wait for ISR signal
// Process all available bytes with preemption enabled
while (uart_rx_ring.tail != uart_rx_ring.head) {
uint8_t b = uart_rx_ring.buf[uart_rx_ring.tail];
uart_rx_ring.tail = (uart_rx_ring.tail + 1) % RING_SIZE;
parse_byte(b); // Safe to take time here
}
}
}
Notice the ISR never disables interrupts, never loops unboundedly (bounded by FIFO depth), and uses vTaskNotifyGiveFromISR which is faster than a queue when you just need a signal. For bare-metal systems, replace the task notify with a volatile flag and check it in your main loop, but ensure the flag is set atomically and the buffer indexes are updated in the right order. A memory barrier (__DMB()) may be needed if the compiler or core can reorder writes.
Protecting Shared State Between ISR and Main Context Without Breaking Latency
Any variable touched by both ISR and thread context is a concurrency bug waiting to happen. The classic mistake is assuming volatile makes access safe. It does not. volatile only tells the compiler not to optimize away loads and stores; it does not make a 32-bit read-modify-write atomic on a 16-bit core, nor does it prevent a task from being preempted mid-update while an ISR fires.
I've found that developers overuse global interrupt disabling to protect shared data. Disabling interrupts for a multi-step update is correct but widens your worst-case latency. A better approach is to minimize shared state and use hardware or compiler guarantees where possible.
Three Proven Patterns for ISR-Task Communication
First, for single-producer single-consumer queues like the ring buffer above, if head is only written by the ISR and tail only by the task, and both are naturally atomic (e.g., 16-bit on Cortex-M3/M4), you do not need to disable interrupts at all. Just ensure the compiler does not tear the read; use volatile and keep the width aligned.
Second, for multi-byte structures, use a lock-free double buffer. The ISR writes to a background buffer, then atomically swaps a pointer or flag indicating which buffer is ready. The task copies the ready buffer with a brief critical section.
Third, when you must share a counter or flag that is modified in both contexts, use atomic access or a brief critical section with BASEPRI masking, not global disable. On Cortex-M4/M7, LDREX/STREX can help, but simply disabling interrupts at the task level for 5-10 cycles is often the most deterministic.
// Pattern: atomic flag swap with minimal masking
// ISR writes sensor sample; task reads it without race
typedef struct { int16_t x; int16_t y; int16_t z; } sample_t;
static volatile sample_t isr_sample;
static volatile uint32_t sample_ready = 0;
void ADC_IRQHandler(void)
{
// Read hardware atomically, clear flag
isr_sample.x = (int16_t)ADC1->DR;
// ... read other channels ...
sample_ready = 1; // Single-word write is atomic on Cortex-M
// No need to disable interrupts inside ISR
}
bool get_latest_sample(sample_t *out)
{
bool has_data = false;
// Briefly mask interrupts up to SYSCALL priority, not globally
UBaseType_t mask = taskENTER_CRITICAL_FROM_ISR(); // or __get_BASEPRI / __set_BASEPRI
// In bare-metal: uint32_t primask = __get_PRIMASK(); __disable_irq();
if (sample_ready) {
*out = (sample_t)isr_sample; // struct copy while masked
sample_ready = 0;
has_data = true;
}
taskEXIT_CRITICAL_FROM_ISR(mask);
// In bare-metal: __set_PRIMASK(primask);
return has_data;
}
Be careful with non-atomic structures. On Cortex-M0, a 32-bit access is atomic, but a 64-bit timestamp is not. If your ISR increments a 64-bit microsecond counter, a task reading it can see a torn value (low word updated, high word not yet). In that case, read in a loop until two consecutive high words match, or protect with a critical section. This class of bug is nearly impossible to catch without deliberate testing, which is why techniques from Embedded Debugging with JTAG and SWD: Tools, Techniques and Workflows like watchpoints on shared variables and non-intrusive trace are invaluable.
For dynamic allocation inside ISRs, don't. Ever. I link to Memory Management in Embedded C: Static Allocation, Pool and Arena Patterns frequently because the same reasoning applies: ISRs must use statically allocated pools or fixed-size buffers. Calling malloc from an ISR is non-deterministic and not reentrant on most libcs, and FreeRTOS pvPortMalloc is explicitly forbidden from ISR context.
Spotting and Fixing Priority Inversion, Starvation, and Spurious Interrupts
Not all interrupt bugs are crashes. Some are subtle performance degradations. Priority inversion in an interrupt context looks like this: a low-priority ISR holds a resource (or masks a higher-priority interrupt implicitly by running) and prevents a high-priority ISR from meeting its deadline. On bare metal, this happens when you spend too long in any ISR; while that ISR runs, all lower or equal priority interrupts are blocked. On RTOS, it happens when a medium-priority task preempts a low-priority task that holds a mutex needed by a high-priority task that was woken by an ISR — the classic inversion, but triggered by interrupt-driven events.
Interrupt starvation is the opposite: a high-rate interrupt continuously preempts a lower-priority one, so the lower one never completes. I've seen this with a misconfigured DMA half-transfer interrupt firing at 48kHz while a 1kHz control timer was set at a lower preemption level. The control loop jittered by hundreds of microseconds. The fix was not to optimize the DMA handler but to lower its preemption priority below the control loop, and to throttle its rate — batch 16 samples per interrupt instead of one.
Spurious interrupts deserve mention. On shared interrupt vectors (e.g., EXTI9_5 covering lines 5-9), you must check which line actually fired. Failing to do so and blindly clearing flags can mask a pending interrupt or clear the wrong one, leaving the line asserted and re-entering the ISR infinitely. Always read the pending register, handle only asserted sources, and verify your clear operation matches the hardware requirement (write-1-to-clear vs read-to-clear).
| Deferral Mechanism | Typical Latency to Handler | Can Block / Allocate? | Best Use Case |
|---|---|---|---|
| Direct in-ISR Processing | Fastest (0 µs deferral) | Never | Only for <2 µs hardware latching (capture, fault pin) |
| Volatile Flag + Main Loop Poll | Polling interval + jitter | No, poll only | Ultra-simple bare-metal, low-rate events (button, slow ADC) |
| Ring Buffer + Task Notification | <1 µs + task switch (2-5 µs) | Deferred task can | High-rate streaming (UART, SPI, sensor DMA) |
| Queue (xQueueSendFromISR) | Task switch + copy overhead | Deferred task can | Discrete messages needing ordering (commands, CAN frames) |
| Semaphore / Event Group | Task switch | Deferred task can | Single event signaling without data payload |
Use this table when choosing your ISR exit strategy. In my experience, task notifications and lock-free ring buffers cover 80% of high-performance cases with the lowest overhead. Queues are better when you need to pass structured messages and preserve ordering, but they copy data and cost cycles inside the ISR. Flags are fine for bare-metal superloops but introduce polling latency that breaks determinism if your main loop has variable execution time.
Quantifying Real-World Latency: From Cycle Counting to Logic Analyzer Traces
You cannot optimize what you do not measure. Datasheet latency numbers are best-case with zero wait states and no contention. I validate every time-critical interrupt with hardware measurement, not simulation alone.
The cheapest method is GPIO toggling. Set a pin high as the first instruction in your ISR and low as the last, and trigger the interrupt with another pin or a PWM edge you control. A 100MHz logic analyzer then shows you interrupt entry latency, ISR duration, preemption nesting, and jitter across thousands of iterations. For more precise CPU-cycle measurement, use the DWT cycle counter. It runs at core clock and can timestamp event entry with single-cycle resolution, though you must enable DWT_CTRL.CYCCNTENA and handle overflow.
Building a Latency Test Harness
For RTOS systems, I inject worst-case load while measuring. Have a low-priority task continuously enter critical sections, allocate memory, or flood the bus with DMA, while your high-priority interrupt fires at its maximum rate. Record the longest observed latency over minutes, not seconds. I also use ARM's exception trace (ETM/ITM) via SWO when available — it logs exception entry/exit without GPIO overhead, and tools like SEGGER SystemView can correlate ISR entry with task scheduling. If your MCU has no ETM, the simple GPIO method plus a timer capture is still remarkably effective.
Set an assertion for ISR duration. At the end of each critical ISR, check DWT_CYCCNT delta and if it exceeds your budget, log or trigger a fault. This catches regressions early when someone later adds a debug print inside the handler. In one product, this catch saved us: a later driver update added a redundant register read inside the ADC ISR that pushed duration from 1.1µs to 2.8µs, violating the 2µs control deadline, and the assertion flagged it in CI before it shipped.
Finally, document your priority map and latency budgets alongside your schematic. When the next engineer adds a new peripheral, they need to see at a glance where it fits without starving existing handlers. A one-page table listing each interrupt, its preemption priority, worst-case duration, max rate, and whether it uses RTOS APIs has prevented more integration bugs in my projects than any amount of code review.
Frequently Asked Questions
How short should an ISR really be?
As short as necessary to avoid losing data or missing the next interrupt. Aim for under 2-5 microseconds on Cortex-M at 100MHz+ for most peripherals, and under 1 microsecond for hard real-time control loops. If your ISR does more than clear the flag, copy a register to RAM, and signal a task, it is probably too long. Measure duration with GPIO or DWT and defer parsing, filtering, and decision logic.
When is it safe to call FreeRTOS functions from an ISR?
Only when the interrupt's priority is at or below (numerically greater or equal to) configMAX_SYSCALL_INTERRUPT_PRIORITY, and you use the FromISR variants like xQueueSendFromISR or vTaskNotifyGiveFromISR. Interrupts above that threshold are never masked by the kernel and must not call any FreeRTOS API. Check your FreeRTOSConfig.h and verify with configASSERT that your NVIC settings match. See the FreeRTOS Documentation for the exact masking rules for your port.
Should I use interrupt priority 0 for the highest urgency interrupt?
You can, but remember priority 0 is never masked by BASEPRI and cannot use RTOS APIs. It also cannot be preempted by anything. Reserve 0 for true fault handlers like hard-fault-adjacent safety shutdowns or precision capture. For most RTOS-aware fast interrupts, use the highest priority that is still maskable by the kernel (typically 5 on a 4-bit system with configLIBRARY_MAX_SYSCALL_INTERRUPT_PRIORITY=5). This gives you low latency while keeping kernel critical sections intact.
How do I handle shared data between an ISR and a task without disabling all interrupts?
Prefer designs that avoid sharing: use a ring buffer where head is ISR-only and tail is task-only, or a double buffer with an atomic ready flag. When sharing is unavoidable, protect the task-side access with a brief BASEPRI-based critical section (taskENTER_CRITICAL) that only masks up to the RTOS syscall priority, not a global __disable_irq. Keep the protected window to a few instructions — copy the data and exit. On Cortex-M, single aligned 32-bit reads and writes are atomic and may not need masking if you structure the protocol to tolerate stale reads.