RTOS Synchronization: Mutexes, Semaphores and Event Groups

Synchronization is where most intermediate RTOS developers hit their first hard fault that isn't a hard fault. Your tasks run fine in isolation, your logic looks correct on paper, and then you get corrupted sensor data, a mysteriously frozen high-priority task, or an I2C bus that locks up once every three days on the bench. In my experience, nearly all of these failures trace back to the same root cause: two or more contexts accessing a shared resource without a well-chosen synchronization primitive. Getting this right isn't about memorizing API calls — it's about understanding what each primitive guarantees about ownership, blocking, and wake-up conditions. This article breaks down the three workhorses I rely on in almost every FreeRTOS and Zephyr project: mutexes, semaphores, and event groups, with the practical trade-offs and failure modes that only show up under real preemptive load.

The Shared Resource Problem That Synchronization Actually Solves

On bare-metal, you can often get away with disabling interrupts around a critical section. Under a preemptive RTOS, that strategy collapses quickly. Once you have multiple tasks and ISRs competing for the same SPI bus, the same telemetry struct, or the same heap, you need a way to serialize access without polling and without wasting CPU.

A race condition doesn't require much code to appear. Consider a common pattern: a low-priority sensor task writes a 32-byte measurement struct while a higher-priority telemetry task reads it to publish over MQTT. On a 32-bit Cortex-M4, the struct copy is not atomic. If the scheduler preempts the writer halfway through, the reader gets a torn read — half old data, half new. The checksum might still pass if you only checksum individual fields. I've chased this exact bug on a cold-chain logger where temperature spikes of 5 degrees appeared only when Wi-Fi reconnect tasks preempted the sensor filter.

What Counts as a Shared Resource

Anything accessed from more than one task or from both task and ISR context is shared. That includes the obvious peripherals (UART, I2C, SPI, flash), but also the less obvious: global state structs, message queues you peek at, static buffers used by driver libraries, and even the C library's malloc if you haven't configured a thread-safe heap. This is why RTOS Memory Management: Static Allocation, Pools and Heap Strategies matters before you tune synchronization — if your heap isn't protected, no mutex around your application data will save you from corruption underneath.

Blocking vs. Spinning

In an RTOS, synchronization primitives block instead of spin. When a task cannot acquire a mutex, it enters the Blocked state and the scheduler runs something else. This is fundamentally different from a spinlock on multicore Linux. Blocking is what lets you build low-power, real-time systems: a task waiting for a sensor interrupt can sleep at 2uA instead of burning cycles. But it also means you must think about blocking time, timeout values, and what should happen when acquisition fails. I've found that hard-coding portMAX_DELAY everywhere is a fast path to a watchdog reset that tells you nothing.

Mutexes: Ownership, Priority Inheritance and When to Protect Data

A mutex is the right tool when you need exclusive, ownership-based access to a resource. The key word is ownership. Only the task that takes a mutex may give it back. That constraint is what enables the RTOS to solve priority inversion for you.

Use a mutex when the question is "who is inside the critical section right now?" If you are protecting data — a shared struct, a peripheral, a file system — you almost always want a mutex, not a semaphore. In FreeRTOS, a mutex is actually a special type of binary semaphore with priority inheritance built in, but you should still create it with xSemaphoreCreateMutex() or, for determinism, xSemaphoreCreateMutexStatic() so you don't fragment the heap at runtime.

Priority Inheritance in Practice

Priority inversion is the classic failure: a low-priority task holds a mutex needed by a high-priority task, but a medium-priority task preempts the low-priority holder and the high-priority task starves indefinitely. I ran into this on a motor controller where a low-priority logging task held the SPI mutex, a medium-priority LED task kept the CPU busy, and the high-priority current-control loop missed its 1 kHz deadline. The symptom was audible jitter.

FreeRTOS mutexes enable priority inheritance by default. When a high-priority task blocks on a mutex held by a lower-priority task, the holder is temporarily boosted to the waiter's priority. Zephyr's k_mutex behaves the same way. This doesn't prevent inversion — it bounds it. If you need stricter guarantees, look at priority ceiling (Zephyr's k_mutex supports it via configuration) or redesign to shorten the critical section. For a refresher on how priorities and preemption interact before you tune inheritance, FreeRTOS Task Scheduling: Priorities, Preemption and Time Slicing is a solid foundation.

Implementing a Mutex-Protected Peripheral

Keep critical sections short and never call blocking functions while holding a mutex. In particular, never do a blocking queue send or a delay inside a mutex hold — you will hold up every other task that needs that resource and create a convoy effect.

// FreeRTOS: Static mutex protecting a shared I2C bus
static SemaphoreHandle_t xI2CMutex;
static StaticSemaphore_t xMutexBuffer;

void vI2CInit(void) {
    xI2CMutex = xSemaphoreCreateMutexStatic(&xMutexBuffer);
}

bool bSensorRead(uint8_t addr, uint8_t *data, size_t len) {
    // Use a finite timeout - don't hang forever on a dead bus
    if (xSemaphoreTake(xI2CMutex, pdMS_TO_TICKS(50)) != pdTRUE) {
        return false; // bus busy or holder crashed
    }

    // --- Critical section: only I2C transactions here ---
    bool ok = i2c_transfer(addr, data, len);
    // --- End critical section ---

    xSemaphoreGive(xI2CMutex);
    return ok;
}

// Recursive mutex variant for nested driver calls
// Only use if you truly need re-entrancy; it hides design flaws
// xSemaphoreCreateRecursiveMutexStatic() + xSemaphoreTakeRecursive()

I've learned to always use timeouts and to log or count mutex timeouts as a health metric. A single timeout in a week is early warning of a task that is holding the bus too long, often because someone added a retry loop inside the critical section.

Semaphores Unpacked: Counting Resources and Signaling Events

If a mutex is about ownership of data, a semaphore is about counting and signaling. A semaphore has no owner; any task or ISR can give it, and any task can take it. That makes it ideal for two patterns that show up constantly in embedded systems: resource counting and ISR-to-task signaling.

Binary Semaphores for ISR-to-Task Signaling

This is the most common correct use of a binary semaphore. An ISR gives the semaphore to unblock a task that does the real work. You never do heavy processing in the ISR itself — you just signal and exit. FreeRTOS requires the FromISR variant and proper handling of the context-switch request.

// FreeRTOS: Binary semaphore signaled from DMA ISR
static SemaphoreHandle_t xDMADoneSemaphore;
static StaticSemaphore_t xSemBuffer;

void vDMAInit(void) {
    xDMADoneSemaphore = xSemaphoreCreateBinaryStatic(&xSemBuffer);
    // Binary semaphores start empty in FreeRTOS - task will block until first give
}

void DMA1_Channel1_IRQHandler(void) {
    BaseType_t xHigherPriorityTaskWoken = pdFALSE;

    if (DMA_GetITStatus(DMA1_IT_TC1)) {
        DMA_ClearITPendingBit(DMA1_IT_TC1);
        xSemaphoreGiveFromISR(xDMADoneSemaphore, &xHigherPriorityTaskWoken);
        // Request context switch if a higher priority task was unblocked
        portYIELD_FROM_ISR(xHigherPriorityTaskWoken);
    }
}

void vDMATask(void *pvParameters) {
    for (;;) {
        // Block indefinitely until ISR signals completion
        // In production, consider a timeout to detect dead DMA
        if (xSemaphoreTake(xDMADoneSemaphore, pdMS_TO_TICKS(1000)) == pdTRUE) {
            process_dma_buffer();
            start_next_dma_transfer();
        } else {
            handle_dma_timeout(); // hardware may need reset
        }
    }
}

A binary semaphore used this way should be created empty and given only from the ISR. If you find yourself giving it from multiple tasks to signal completion, you probably want an event group or a queue instead.

Counting Semaphores for Resource Pools

A counting semaphore tracks how many instances of a resource are available. I use this for DMA descriptor pools, fixed-size memory blocks, and connection slots. You initialize the count to N, each taker decrements it, and each releaser increments it. When the count hits zero, takers block.

For example, if you have a pool of four reusable telemetry buffers, create a counting semaphore with max 4, initial 4. A task takes the semaphore before acquiring a buffer and gives it when returning the buffer. Unlike a mutex, there is no priority inheritance here — the RTOS can't know which task holds which buffer instance. If priority inversion is a concern with pools, consider partitioning pools by priority or using separate queues per priority class. The FreeRTOS Documentation details the exact create/take/give semantics for counting semaphores and the static allocation options that avoid heap use.

Event Groups and Flags: Orchestrating Multi-Condition Synchronization

Event groups (FreeRTOS) or event flags (Zephyr) solve a different problem: waiting for any combination of conditions. Instead of blocking on a single object, a task can block until bit 0 AND bit 3 are set, or until bit 1 OR bit 2 is set, with optional auto-clear on exit. This is far more efficient than taking multiple semaphores in sequence, which can deadlock, or polling a set of flags in a loop.

I reach for event groups when a task needs to synchronize to system state rather than to a single resource. Typical cases: wait until both Wi-Fi is connected AND flash is initialized before starting an OTA task; sleep until either a button press OR a BLE command arrives; coordinate a startup sequence where three sensor drivers each set their ready bit.

How Event Groups Work Under the Hood

Internally, an event group is a single 32-bit (or 8/24-bit on some ports) integer where each bit is an event flag. Tasks set bits with xEventGroupSetBits() or xEventGroupSetBitsFromISR(), and waiters call xEventGroupWaitBits() specifying a mask, whether all bits are required, and whether to clear on exit. In Zephyr, the equivalent is k_event with k_event_set() and k_event_wait().

The atomicity of the wait is the key advantage. The test-and-clear is performed atomically inside the kernel, so you won't miss a short pulse that occurs between checking and clearing. If you try to replicate this with separate flags and a mutex, you will inevitably introduce a window where an event is lost.

// FreeRTOS: Startup orchestration with event groups
#define WIFI_READY_BIT      (1U << 0)
#define SENSORS_READY_BIT   (1U << 1)
#define FLASH_READY_BIT     (1U << 2)
#define ALL_REQUIRED_BITS   (WIFI_READY_BIT | SENSORS_READY_BIT | FLASH_READY_BIT)

static EventGroupHandle_t xSystemEvents;
static StaticEventGroup_t xEventGroupBuffer;

void vSystemInit(void) {
    xSystemEvents = xEventGroupCreateStatic(&xEventGroupBuffer);
}

// Called from Wi-Fi task when connected
void vWiFiNotifyReady(void) {
    xEventGroupSetBits(xSystemEvents, WIFI_READY_BIT);
}

// Initialization task waits for all subsystems
void vStartupTask(void *pvParameters) {
    // Wait for ALL bits, clear on exit, timeout 10 seconds
    EventBits_t uxBits = xEventGroupWaitBits(
        xSystemEvents,
        ALL_REQUIRED_BITS,
        pdTRUE,          // clear bits after wait
        pdTRUE,          // wait for ALL bits
        pdMS_TO_TICKS(10000)
    );

    if ((uxBits & ALL_REQUIRED_BITS) == ALL_REQUIRED_BITS) {
        start_application_tasks();
    } else {
        // Handle missing subsystem - log which bit never arrived
        handle_startup_failure(uxBits);
    }
}

// Alternative: wait for ANY button or BLE event
// xEventGroupWaitBits(xSystemEvents, BTN_BIT | BLE_CMD_BIT, pdTRUE, pdFALSE, portMAX_DELAY)

Practical Limitations

Event groups are not a queue. If the same bit is set twice before any task waits on it, it still only counts as one event — the second set is coalesced. If you need to count every occurrence, use a counting semaphore or a queue. Also, keep ISR usage lightweight: xEventGroupSetBitsFromISR() on FreeRTOS actually defers work to the timer task on many ports, which adds a few microseconds of latency compared with a direct semaphore give. For hard real-time ISR signaling where latency matters, I still prefer a binary semaphore; the Zephyr Project Documentation has a helpful discussion of these trade-offs for k_event vs k_sem.

Deadlocks, Priority Inversion and Other Synchronization Failures I've Debugged

Most synchronization bugs don't crash immediately. They stall, corrupt data occasionally, or pass all unit tests and fail once a week in the field. Here are the patterns I watch for during code review.

The Circular Wait Deadlock

Deadlock requires four conditions, but the one you can control is circular wait: Task A holds Mutex X and waits for Mutex Y, while Task B holds Y and waits for X. The fix is a global lock ordering rule. Document the order — for example, always acquire I2CMutex before FlashMutex — and enforce it in review. If you need both, acquire them together with a timeout and back off if you can't get both. I've also seen deadlock caused by a task taking the same non-recursive mutex twice due to a refactored call chain; enabling configASSERT on mutex take will catch that in FreeRTOS immediately.

Accidental Priority Inversion Without a Mutex

Using a binary semaphore as a mutex is the most common way to accidentally disable priority inheritance. Because the semaphore has no owner, the kernel can't boost the holder. I've seen this in code that used xSemaphoreCreateBinary() for I2C protection because "it was one line shorter." It worked until system load increased, then high-priority control tasks started missing deadlines by 40 ms. The fix was a one-line change to xSemaphoreCreateMutex(), but the debug took two days with a logic analyzer. If you need Real-Time Debugging: Trace Tools, Logic Analyzers and JTAG Techniques, this is exactly the scenario where a trace tool like Percepio Tracealyzer or SEGGER SystemView would have shown the inversion in minutes.

Holding Locks Too Long and ISR Misuse

Two other frequent mistakes: doing flash erase or NVS writes while holding a shared bus mutex, and calling non-ISR-safe APIs from an ISR. Flash operations can block for 20-80 ms; holding a mutex across that stalls everything. Split the critical section: copy data under mutex, release, then do the slow operation. And never call xSemaphoreTake() from an ISR — on FreeRTOS it will assert if configured, on Zephyr it will return an error. Only the FromISR / ISR-safe variants are allowed. In my current projects I run a static analysis rule that flags any Take or Give inside a function with _ISR or IRQHandler in the name.

Choosing the Right Primitive: Latency, Footprint and Real-World Tradeoffs

There is no single best primitive — only the one that matches your synchronization semantics with the least coupling. I choose based on three questions: Do I need ownership? Do I need to count events? Do I need to wait for multiple conditions?

For protecting shared data or a peripheral, always start with a mutex. For signaling a single event from ISR to task, use a binary semaphore. For managing a pool of N identical resources, use a counting semaphore. For coordinating startup, modes, or multi-source wake-ups, use an event group.

The table below is the cheat sheet I keep in our team's embedded style guide. It assumes FreeRTOS on Cortex-M, but the concepts map directly to Zephyr's k_mutex, k_sem, and k_event.

Primitive Ownership Primary Use Case Can Be Given from ISR? Priority Inheritance?
Mutex Yes — only holder can give Protect shared data & peripherals No Yes (inheritance / ceiling)
Binary Semaphore No ISR-to-task signaling, single event Yes (FromISR) No
Counting Semaphore No Pool of N resources, throttling Yes (FromISR) No
Event Group / Flags No Wait for AND/OR of multiple conditions Yes (deferred via timer task) No

Beyond semantics, consider footprint and timing. A mutex and a binary semaphore are similar in RAM (roughly 80 bytes for the control block on FreeRTOS) and take/give latency is typically 1-3 microseconds on a 80 MHz Cortex-M4. An event group is cheaper if you replace three semaphores with one 32-bit flag set, both in RAM and in wake-up latency — one wait instead of three sequential waits. On Zephyr, event flags are especially lightweight when configured with CONFIG_EVENTS=y.

One final habit that has saved field deployments: pair every synchronization primitive with a watchdog check. If a high-priority task blocks on a mutex longer than its worst-case expected hold time, let it time out, log the holder, and recover. A task that blocks forever on portMAX_DELAY is indistinguishable from a dead task to the scheduler, and your Watchdog Timers: Building Reliable Embedded Systems That Self-Recover strategy needs to account for that. In several of our products we use a dedicated monitor task that periodically checks uxSemaphoreGetCount() or event bits and petting logic that only succeeds when synchronization health is nominal. It turns silent stalls into actionable resets with diagnostics.

Frequently Asked Questions

Can I use a binary semaphore instead of a mutex to protect shared data?

Technically you can, but you shouldn't. A mutex provides ownership and priority inheritance; a binary semaphore does not. If a low-priority task holds a binary semaphore used as a lock and a medium-priority task preempts it, a high-priority waiter will starve without any boost to the holder. I've debugged field systems where this mistake added 50 ms of jitter to a control loop. Use xSemaphoreCreateMutex() for data protection and reserve binary semaphores for ISR signaling where inheritance isn't applicable.

When should I choose an event group over multiple semaphores?

Choose an event group when you need to wait for a combination of conditions atomically — for example, "wait until both network and sensors are ready" or "wake when either a button or BLE command arrives." Waiting on multiple semaphores sequentially can deadlock and misses the atomic test-and-clear guarantee. Event groups coalesce multiple events into one 32-bit wait, which is both faster and less error-prone. If you need to count every single occurrence of an event, however, stay with a counting semaphore or queue — event bits don't queue.

Is it safe to take a mutex inside an ISR?

No. No RTOS I use allows blocking calls from ISR context, and mutex take can block. In FreeRTOS you must use xSemaphoreGiveFromISR() and xEventGroupSetBitsFromISR() for ISR signaling, and never attempt to take. In Zephyr, k_mutex_lock() from ISR will return an error. Keep ISRs short: clear the interrupt, give a semaphore or set an event bit, and do the mutex-protected work in a high-priority task that unblocks from that signal.

How do I debug a deadlock that only happens intermittently?

Start by enabling kernel diagnostics: FreeRTOS configASSERT, stack overflow checking, and mutex holder tracking. Instrument each mutex take with a finite timeout and log the holder's task handle when a timeout occurs. Then capture a trace with Percepio Tracealyzer or SEGGER SystemView — the visualization will show circular wait and priority inversion immediately. In my experience, enforcing a global lock ordering rule (always take locks in the same documented order) eliminates most deadlocks before they ship. Adding a monitor task that checks blocking durations and feeds a windowed watchdog gives you a recovery path in the field.

Related Articles

References & Standards: FreeRTOS Documentation · Zephyr Project Documentation · MQTT Specification