RTOS Memory Management: Static Allocation, Pools and Heap Strategies

When your RTOS application hard-faults at 2 AM because a sensor task couldn't allocate 64 bytes, you learn quickly that memory management on microcontrollers is not just a software detail — it's a system architecture decision. Unlike Linux where you can rely on virtual memory and an MMU to paper over leaks and fragmentation, an RTOS running on a Cortex-M4 with 128KB of SRAM gives you no safety net. I've spent the last decade shipping firmware where every byte of RAM was accounted for before the scheduler ever started, and the strategies that work are fundamentally different from desktop development. This article covers how I approach rtos memory design: when to lock everything down with static allocation, how to use a memory pool for deterministic behavior, and which heap strategies actually make sense in a real-time context.

Why Dynamic Allocation Breaks Determinism in Real-Time Systems

In my experience, the first mistake teams make is treating an RTOS like a small Linux. They call malloc() inside tasks, allocate message buffers on demand, and assume the heap will sort itself out. It won't. On a typical MCU without an MMU, the heap is just a contiguous block of SRAM managed by a simple linked list or first-fit algorithm. Every allocation takes a variable amount of time, and every free operation can fragment that block into unusable slivers.

Real-time means predictable. If your high-priority motor control task needs to allocate memory and the heap is fragmented, that allocation might take 5 microseconds one time and 500 microseconds the next while the allocator walks a fragmented free list. That jitter is enough to miss a control deadline. Worse, fragmentation is cumulative and non-deterministic — you can't reliably test for it. A system that runs fine for three days in the lab can fail after three weeks in the field.

The Hidden Cost of Fragmentation

Fragmentation happens when you allocate and free blocks of different sizes. Even if you have 20KB free in total, it might be split into 40 scattered 512-byte holes. A request for a 2KB buffer will fail despite sufficient total free memory. First-fit and best-fit allocators try to mitigate this, but they only delay the inevitable. I've debugged systems where heap fragmentation caused a slow memory leak that looked like a software bug, but was really just the allocator's inability to coalesce adjacent free blocks efficiently.

Latency and Failure Modes You Can't Afford

Standard malloc() is also non-deterministic in its execution time and its failure mode. It may need to disable interrupts or take a scheduler lock while it manipulates the free list, blocking other tasks. And when it fails, it returns NULL — which too much application code neglects to check. In a safety-critical loop, an unchecked NULL dereference is a hard fault. For anything with real-time constraints, that combination of unbounded latency and silent failure is unacceptable. This is why many safety standards like MISRA C and certifications for IEC 61508 strongly discourage or forbid heap use after initialization.

Static Allocation: Designing for Predictability Before the Scheduler Starts

Static allocation means every RTOS object — tasks, stacks, queues, semaphores — is allocated at compile time or at startup before the scheduler runs, and never freed. Your RAM usage becomes a known constant that you can verify in the map file. There are no runtime allocation failures because there are no runtime allocations.

This approach forces you to think through your worst-case resource needs upfront. How many messages can be in flight at once? What is the maximum stack depth for each task? You answer these questions during design, not when the device is in the customer's hands. I now default to static allocation for all long-lived objects, even on projects where heap is technically available. The upfront planning pays for itself in debug time saved.

Both FreeRTOS and Zephyr support full static allocation. In FreeRTOS, you enable configSUPPORT_STATIC_ALLOCATION and use the ...Static variants of the creation functions. Instead of the kernel calling pvPortMalloc() internally, you provide the memory.

/* FreeRTOS: Creating a task with fully static allocation */
#define TASK_STACK_SIZE 256

static StackType_t xTaskStack[TASK_STACK_SIZE];
static StaticTask_t xTaskBuffer;
static TaskHandle_t xHandle = NULL;

void create_sensor_task(void)
{
    xHandle = xTaskCreateStatic(
        vSensorTask,          /* Task function */
        "Sensor",             /* Name for debugging */
        TASK_STACK_SIZE,      /* Stack depth in words, not bytes */
        NULL,                 /* Task parameter */
        2,                    /* Priority - see note on scheduling */
        xTaskStack,           /* Stack buffer provided by application */
        &xTaskBuffer);        /* TCB buffer provided by application */

    /* xHandle will be NULL only if parameters are invalid.
       No heap involved, so no allocation failure due to fragmentation. */
}

/* Static queue creation - no heap */
static uint8_t ucQueueStorage[10 * sizeof(SensorReading_t)];
static StaticQueue_t xQueueBuffer;
static QueueHandle_t xSensorQueue;

void create_sensor_queue(void)
{
    xSensorQueue = xQueueCreateStatic(
        10,                   /* Queue length */
        sizeof(SensorReading_t),
        ucQueueStorage,
        &xQueueBuffer);
}

Notice that stack size is specified in words (4 bytes on Cortex-M). A common error I see is allocating 256 bytes when you meant 256 words, leading to immediate stack overflow. Your linker script should place xTaskStack and ucQueueStorage in SRAM, and you can confirm their exact addresses and sizes in the .map file.

When Static Allocation Is Non-Negotiable

Use static allocation for anything that lives for the lifetime of the system: task control blocks (TCBs) and stacks, inter-task communication primitives, and driver buffers. If you are working with RTOS Synchronization: Mutexes, Semaphores and Event Groups, those primitives should also be statically allocated. A mutex that protects a shared I2C bus should not be dynamically created when a task first needs it. I've found that requiring all such objects to be static makes system reviews far easier — you can grep the codebase for any remaining calls to xTaskCreate() or xQueueCreate() and eliminate them.

Zephyr takes this further with its device tree and Kconfig model. Many objects are defined at build time via macros like K_THREAD_DEFINE and K_MSGQ_DEFINE, which place the thread stack and queue buffer in dedicated sections. If you are coming from FreeRTOS, the Getting Started with Zephyr RTOS: Device Tree, Kconfig and Your First Application guide is a good primer on how that build-time model changes your memory layout thinking.

Memory Pools: Fixed-Block Allocation Without Fragmentation

Static allocation is perfect for fixed resources, but what about data that comes and goes — like network packets, sensor samples, or command messages? You don't want to statically allocate a buffer for every possible packet if you might have hundreds per second, but you also can't use malloc() for them. This is where a memory pool (also called a memory slab or fixed-size block pool) comes in.

A memory pool is a pre-allocated array of identical blocks. Allocation is just taking the next free block off a linked list; freeing is returning it. Both are O(1) and deterministic, typically just a few instructions with interrupts masked for a very short window. Because all blocks are the same size, there is zero external fragmentation. A pool of 32 blocks of 128 bytes will always be able to satisfy 32 allocations of 128 bytes, no matter what order you allocate and free them.

In my designs, I usually create one pool per message type or packet size. A telemetry system I worked on had three pools: 64-byte blocks for sensor readings, 256-byte blocks for BLE packets, and 1024-byte blocks for log messages. This avoids internal fragmentation (wasting space by putting a 60-byte message in a 256-byte block) while keeping each pool simple.

/* Zephyr: Using k_mem_slab for deterministic packet allocation */
K_MEM_SLAB_DEFINE(packet_slab, 128, 32, 4);

struct packet {
    sys_snode_t node;
    uint8_t data[128 - sizeof(sys_snode_t)];
};

void producer_thread(void)
{
    struct packet *pkt;
    int ret;

    /* K_NO_WAIT = non-blocking, deterministic.
       K_FOREVER would block and affect schedulability. */
    ret = k_mem_slab_alloc(&packet_slab, (void **)&pkt, K_NO_WAIT);
    if (ret == 0) {
        read_sensor(pkt->data);
        k_msgq_put(&packet_queue, &pkt, K_NO_WAIT);
    } else {
        /* Pool exhausted - count and handle, don't crash.
           This is a backpressure signal. */
        atomic_inc(&dropped_packets);
    }
}

void consumer_thread(void)
{
    struct packet *pkt;
    k_msgq_get(&packet_queue, &pkt, K_FOREVER);
    process_and_transmit(pkt);
    k_mem_slab_free(&packet_slab, (void **)&pkt);
}

/* FreeRTOS equivalent using heap_4 with fixed-size allocation
   wrapper is less efficient - prefer Stream Buffers or
   statically allocated pools with FreeRTOS Memory Pools (from CMSIS-RTOS)
   if available. */

Sizing Pools and Handling Exhaustion

The hardest part is sizing the pool correctly. You need to handle the worst-case burst, not the average throughput. I calculate this by looking at the producer rate, consumer rate, and maximum blocking time of the consumer. If your producer can generate 100 messages per second and the consumer might be blocked for 200ms (perhaps waiting on flash), you need at least 20 blocks just to buffer that window, plus margin. Add 25-30% headroom and then instrument the high-water mark.

Exhaustion must be a handled case, not a fatal error. In the code above, we increment a counter and drop the packet. For non-critical telemetry that is correct. For critical commands, you might instead block with a timeout or signal backpressure to the producer. What you must never do is spin waiting for a free block in a high-priority task — that can deadlock the system if the low-priority consumer is the one that would free the block. This interaction between memory pools and FreeRTOS Task Scheduling: Priorities, Preemption and Time Slicing is critical to get right.

Heap Strategies in RTOS: Heap_1 Through Heap_5 and When to Use Them

Sometimes you do need a general-purpose heap — for example, during initialization to parse configuration, or for a third-party library that calls malloc(). FreeRTOS famously provides five heap implementations rather than one, and the choice matters enormously. I've seen projects ship with the wrong heap and wonder why they run out of memory after a week. The FreeRTOS Documentation details each, but here is how I think about them in practice.

Strategy Alloc / Free Support Determinism Fragmentation Handling Practical Use Case
Heap_1 Alloc only, no free Deterministic None (no free = no fragmentation) Simple systems that allocate once at startup. Safest if you can use it.
Heap_2 Alloc + Free, no coalescence Non-deterministic None - suffers from fragmentation Deprecated. Do not use in new designs.
Heap_3 Wraps standard malloc/free Depends on toolchain Depends on toolchain Only for compatibility with libraries requiring malloc. Requires thread-safety wrappers.
Heap_4 Alloc + Free with coalescence Non-deterministic Coalesces adjacent free blocks General purpose. Good default for most projects that need free(). Single RAM region.
Heap_5 Alloc + Free with coalescence Non-deterministic Coalesces, spans multiple regions Needed when RAM is split (e.g., internal SRAM + external SDRAM).

In my experience, 90% of projects that think they need a heap should use Heap_4 and nothing else. It coalesces adjacent free blocks, which significantly reduces fragmentation compared to Heap_2, and it includes a configAPPLICATION_ALLOCATED_HEAP option so you can place the heap array exactly where you want via the linker. Heap_1 is even safer — if your system allocates all tasks and queues dynamically at boot and never frees them, Heap_1 guarantees you will never fragment and allocations will be fast and predictable. I've used Heap_1 on certified firmware where dynamic allocation after startup was simply forbidden.

Heap_3 is the one to be careful with. It just wraps your compiler's malloc(). That implementation may not be thread-safe, may use semihosts, or may have terrible fragmentation behavior. If you must use Heap_3 because a library demands it, you must provide __malloc_lock and __malloc_unlock or equivalent to make it thread-safe, and you must understand your toolchain's allocator.

Zephyr's Heap Options

Zephyr offers k_heap and k_mem_pool with a TLSF (Two-Level Segregated Fit) allocator that is O(1) and designed for real-time use. According to the Zephyr Project Documentation, TLSF gives bounded allocation time even with fragmentation. It's a better general-purpose heap than FreeRTOS Heap_4 if you need true dynamic behavior. Still, I treat any heap as a convenience for initialization and non-real-time threads. The real-time data path should remain on pools or static buffers.

Guarding Against Stack Overflow and Heap Exhaustion at Runtime

Choosing the right strategy is only half the job; you have to verify it won't fail at runtime. Stack overflow is the most common memory bug I see on RTOS projects. Tasks silently write past their stack into the next TCB or heap block, corrupting the system in ways that manifest minutes later as a mysterious crash in an unrelated task.

Both FreeRTOS and Zephyr provide stack overflow detection, but it is not enabled by default and has limitations. FreeRTOS offers two methods via configCHECK_FOR_STACK_OVERFLOW. Method 1 checks the stack pointer on context switch — cheap but catches overflow late. Method 2 fills the last 16 bytes of the stack with a known pattern (0xA5) and checks if that pattern is overwritten. It catches overflow faster but can miss a large overflow that jumps over the pattern. I enable Method 2 during development and also add explicit watermarking.

/* FreeRTOS: Enabling runtime checks in FreeRTOSConfig.h */
#define configCHECK_FOR_STACK_OVERFLOW  2
#define configUSE_MALLOC_FAILED_HOOK    1
#define configUSE_STATS_FORMATTING_FUNCTIONS 1

/* Application hooks - implement these */
void vApplicationStackOverflowHook(TaskHandle_t xTask, char *pcTaskName)
{
    /* Log task name to persistent storage if possible, then reset.
       Do NOT try to recover - stack is already corrupted. */
    log_error("STACK OVERFLOW in %s", pcTaskName);
    taskENTER_CRITICAL();
    /* Trigger watchdog or system reset */
    NVIC_SystemReset();
}

void vApplicationMallocFailedHook(void)
{
    /* Heap exhausted. This fires for pvPortMalloc failures.
       Log heap high-water mark and current free heap. */
    log_error("MALLOC FAILED free=%u min_free=%u",
              xPortGetFreeHeapSize(),
              xPortGetMinimumEverFreeHeapSize());
    /* For critical systems, reset here. For others, degrade gracefully. */
}

/* Manual stack watermark check - call periodically from monitor task */
void check_stack_highwater(void)
{
    /* uxTaskGetStackHighWaterMark returns words free since task start */
    UBaseType_t watermark = uxTaskGetStackHighWaterMark(xHandle);
    if (watermark < 32) { /* Less than 128 bytes on Cortex-M4 */
        log_warn("Task %s low stack: %u words free", 
                 pcTaskGetName(xHandle), watermark);
    }
}

Heap Monitoring and the Minimum Ever Free Metric

The single most useful heap metric is xPortGetMinimumEverFreeHeapSize(). It tells you the lowest amount of free heap that has existed since boot — your true worst-case. I log this value every hour and after any major operation. If it trends downward over days, you have a leak or fragmentation issue. If it drops suddenly after a new feature, you under-sized the heap.

For pool monitoring, track the number of free blocks. Zephyr's k_mem_slab_num_free_get() and FreeRTOS's uxQueueSpacesAvailable() for queue-based pools give you this. Set a threshold alert at 20% free — if you hit it, you are closer to exhaustion than you think, because bursts are bursty.

On Cortex-M3/M4/M7 parts with an MPU, enable it. You can configure the MPU to make the last 32 bytes of each task stack inaccessible, so any overflow triggers an immediate MemManage fault instead of silent corruption. Combined with Real-Time Debugging: Trace Tools, Logic Analyzers and JTAG Techniques like SEGGER SystemView or Percepio Tracealyzer, you can visualize stack usage and heap fragmentation over time rather than guessing.

Sizing Your Memory Layout: From Linker Script to Runtime Validation

All of this comes together in the linker script. Your RAM is not just "heap" — it's a carefully partitioned layout: static buffers, task stacks, pools, heap, and interrupt stack. I insist on a linker script that explicitly places each region and reserves guard gaps.

A typical layout for a 128KB SRAM device might be: 4KB for interrupt stack and vectors, 40KB for statically allocated task stacks and TCBs, 20KB for memory pools, 16KB for DMA buffers (often in a non-cacheable region), and the remaining 48KB for Heap_4. The key is that none of these regions overlap and you can prove it from the map file. The _estack, _sheap, and related symbols should be defined in the linker script, not guessed in code.

Practical Sizing Rules I Follow

For task stacks, I start with 512 words (2KB) for simple tasks, 1024 words for tasks that call printf or do floating point, and measure the high-water mark under worst-case conditions (all interrupts firing, all branches taken). Then I add 30% margin. Never size a stack based on "it seems to work" — measure it. Tools like uxTaskGetStackHighWaterMark() or Zephyr's k_thread_stack_space_get() exist for this reason. The worst stack overflows I've debugged came from error-handling paths that were rarely executed but called deep formatting functions.

For heap, if you must have one, size it after accounting for everything else. The heap is what is left, not what you start with. During bring-up, fill the entire heap with a pattern (0xAA), run the system through a stress test that exercises every feature — including OTA updates and error recovery — and then check how much of that pattern remains untouched. That untouched area is your margin. If your margin is less than 15% of the heap, you are too tight. I've had products pass all functional tests but fail after six months because a rarely-used diagnostic command allocated a large temporary buffer that fragmented the heap just enough to make the next allocation fail.

Finally, validate at boot. Add a startup self-test that checks that the heap size matches expectations (xPortGetFreeHeapSize() on FreeRTOS should equal your configured heap within a few bytes), that each memory pool has the expected number of free blocks, and that no stack high-water mark is already low. If any check fails, fail to boot with a clear error code rather than limping along. A device that refuses to boot with "heap size mismatch 0x02" is far easier to diagnose than one that crashes randomly three hours later. Pair this with a robust recovery strategy like the one described in Watchdog Timers: Building Reliable Embedded Systems That Self-Recover so that a memory-related reset still leaves the system in a safe state.

Frequently Asked Questions

Can I safely use malloc and free after the scheduler has started?

You can, but you should isolate it. Use Heap_4 or Zephyr's TLSF heap and restrict allocations to low-priority, non-real-time tasks. Never call malloc/free from an ISR or from a high-priority control loop. In my experience, the safest pattern is to allow heap use only during initialization and then switch to pools and static buffers for the steady-state runtime. If you enable configUSE_MALLOC_FAILED_HOOK, always implement the hook to log and handle failures.

How do I choose between a memory pool and a queue for passing data between tasks?

They solve different problems and often work together. A queue copies data; a memory pool passes ownership of a buffer. For small, fixed-size messages (like a sensor reading struct), a queue that copies the data is simpler and avoids pool management. For larger or variable-size data (like network packets or audio frames), copying through a queue wastes CPU and RAM. Instead, allocate a block from a memory pool, fill it, and send its pointer through a queue. The receiver processes the data and returns the block to the pool. This zero-copy pattern is essential for throughput.

My xPortGetMinimumEverFreeHeapSize keeps decreasing slowly. Is that a leak?

It could be a leak, but it is more often fragmentation. A true leak is when you allocate and never free — check that every pvPortMalloc has a matching vPortFree on all code paths, including error paths. Fragmentation shows as a sawtooth pattern: free heap goes down as you allocate, up as you free, but the minimum ever free trends down because free blocks are scattered. If you see this, switch from heap allocations to a fixed-size memory pool for the most frequent allocation size, or move to Heap_4 if you are still on Heap_2.

How large should I make each task stack?

There is no universal size. Start with an estimate (512-1024 words), enable stack overflow detection (configCHECK_FOR_STACK_OVERFLOW = 2) and watermark filling, then run a worst-case stress test that includes interrupts, nested calls, and error handling. Check uxTaskGetStackHighWaterMark for each task. The amount of stack actually used is (allocated size - high-water mark). Add 30% margin to the used amount and round up to the next power of two if your MPU requires alignment. Re-measure after any significant code change to that task.

Related Articles

References & Standards: FreeRTOS Documentation · Zephyr Project Documentation · MQTT Specification