Bare Metal vs RTOS: When Each Approach Makes Sense

Bare metal versus RTOS is one of the first architectural decisions you will make on any microcontroller project, and it is also one of the most consequential to reverse. In my experience, teams often frame the choice as a matter of complexity — bare metal is simple, an RTOS is complex — when in reality it is a trade-off between deterministic control and composable concurrency. I have shipped both: a battery-powered BLE sensor node that ran for two years on a single coin cell with not a single task in sight, and a gateway controller juggling LTE-M, MQTT, and local control loops where trying to do it all in a super-loop would have been unmaintainable. Neither approach is inherently better. The right choice depends on timing requirements, available memory, team experience, and how your firmware needs to grow over the next 18 months. This article breaks down what each approach actually entails on modern ARM Cortex-M hardware, where the hidden costs sit, and how I decide which path to take for IoT products.

What "Bare Metal" Really Looks Like Inside the Super-Loop

When we say bare metal, we mean firmware that runs directly on hardware without a kernel scheduling threads for you. There is no scheduler, no task control blocks, and no OS-provided synchronization primitives. Your entry point is main(), and after initializing clocks, GPIOs, and peripherals, you typically fall into an infinite loop — the super-loop — that polls flags, handles state machines, and sleeps when there is nothing to do.

That simplicity is deceptive. You still have concurrency, it just comes from interrupts. A bare metal system is really two domains: the main loop (background) and interrupt service routines (foreground). Any real-time work must either happen inside an ISR or be deferred by setting a flag the super-loop will act on. I have found that this model works exceptionally well when your product has one primary job and a handful of well-understood peripheral events.

The Anatomy of a Super-Loop With Sleep

On a Cortex-M0+ or M4, a clean bare metal main loop is not just while(1) spinning at full speed. You use __WFI() (Wait For Interrupt) to let the core sleep until an event arrives, which is critical for low-power IoT devices. Interrupts wake the core, set volatile flags or fill a ring buffer, and return. The loop then processes those flags sequentially.

#include <stdbool.h>
#include <stdint.h>
#include "stm32l4xx.h"

volatile bool adc_data_ready = false;
volatile uint16_t adc_sample = 0;

// Minimal ISR: do as little as possible
void ADC1_IRQHandler(void) {
    if (ADC1->ISR & ADC_ISR_EOC) {
        adc_sample = (uint16_t)ADC1->DR;
        adc_data_ready = true;
        ADC1->ISR |= ADC_ISR_EOC; // clear flag
    }
}

int main(void) {
    SystemClock_Config();
    GPIO_Init();
    ADC_Init_InterruptMode();
    __enable_irq();

    while (1) {
        // Sleep until an interrupt fires - current drops to microamps
        __WFI();

        if (adc_data_ready) {
            adc_data_ready = false;
            // Process outside ISR: filtering, threshold check
            int16_t filtered = process_sample(adc_sample);
            if (filtered > THRESHOLD) {
                trigger_actuator();
            }
            send_ble_notification(filtered);
        }
        handle_button_state_machine();
    }
}

This pattern keeps ISR latency in the low microseconds and avoids priority inversions entirely because there is no blocking. But it forces you to think in state machines. If you need to wait 50ms for a sensor to stabilize after power-on while also debouncing a button and servicing UART, you cannot call delay(50). You need non-blocking timers, usually driven by SysTick or a hardware timer, and you must manually slice your logic so no single iteration starves the rest. For two or three concurrent activities, this is manageable. Beyond that, the super-loop becomes a tangle of flags and counters that is hard to test and easy to break, which is why understanding Interrupt Handling Best Practices: Priority, Latency and ISR Design is essential even before you consider an RTOS.

Where Bare Metal Shines in Practice

In my experience, bare metal is the clear winner in three situations. First, ultra-low-power sensors where every byte of RAM and every microamp matters — think a LoRa soil moisture node on an STM32L0 with 20KB RAM. Second, safety or certification-constrained code where you need full traceability and want to avoid certifying a kernel. Third, very fast, jitter-sensitive control loops (motor control, high-speed sampling) where you need to guarantee a 10 microsecond ISR response without any scheduler overhead. The code is also easier to bring up: no linker script modifications for heap, no kernel port, and debugging is straightforward because the call stack is always just main -> function or an ISR.

How an RTOS Kernel Earns Its Keep: Tasks, Scheduling and Inter-Task Communication

An RTOS does not make your microcontroller faster. It makes concurrency explicit and manageable by giving you threads (called tasks in FreeRTOS and Zephyr) that the kernel schedules. On Cortex-M, this typically means a preemptive priority-based scheduler driven by the SysTick interrupt. The highest-priority ready task runs; if it blocks or a higher-priority task becomes ready, the kernel performs a context switch — saving registers to the task stack and restoring the next task's context via PendSV.

What you gain is the ability to write blocking code without blocking the whole system. A task can call vTaskDelay(100) or wait on a queue, and the kernel will automatically run other tasks while it waits. This transforms architecture from one giant state machine into several small, sequential programs that interact through well-defined primitives: queues, semaphores, mutexes, and event groups.

Tasks and Queues Instead of Flags and Super-Loops

Consider an IoT thermostat that reads a temperature sensor over I2C, runs a PID loop, handles a Wi-Fi module over UART, and publishes via MQTT. In bare metal you would interleave all of that. With an RTOS, each concern becomes a task:

// FreeRTOS example: sensor task blocks on delay, comms task blocks on queue
#include "FreeRTOS.h"
#include "task.h"
#include "queue.h"
#include "semphr.h"

QueueHandle_t sensorQueue;
SemaphoreHandle_t i2cMutex;

void vSensorTask(void *pvParameters) {
    SensorReading reading;
    for (;;) {
        // Take mutex so Wi-Fi task doesn't collide on shared I2C bus
        if (xSemaphoreTake(i2cMutex, pdMS_TO_TICKS(100)) == pdTRUE) {
            reading = read_sht31_blocking(); // can block 15ms, that's ok
            xSemaphoreGive(i2cMutex);
            xQueueSend(sensorQueue, &reading, 0);
        }
        vTaskDelay(pdMS_TO_TICKS(1000)); // sleep 1s, other tasks run
    }
}

void vMqttTask(void *pvParameters) {
    SensorReading rx;
    for (;;) {
        // Block indefinitely until data arrives - no polling needed
        if (xQueueReceive(sensorQueue, &rx, portMAX_DELAY) == pdTRUE) {
            publish_mqtt_temperature(rx.temperature, rx.humidity);
        }
    }
}

int main(void) {
    hardware_init();
    sensorQueue = xQueueCreate(8, sizeof(SensorReading));
    i2cMutex = xSemaphoreCreateMutex();
    xTaskCreate(vSensorTask, "SENSOR", 256, NULL, 2, NULL);
    xTaskCreate(vMqttTask, "MQTT", 512, NULL, 1, NULL);
    vTaskStartScheduler(); // never returns
    for(;;);
}

This code is longer, but each task reads like a simple sequential program. The queue provides thread-safe buffering and backpressure, and the mutex prevents bus contention without you building your own lock. The FreeRTOS Documentation describes task priorities, stack sizing, and blocking semantics in detail, and I recommend reading it before sizing your first task — undersized stacks are the single most common RTOS crash I see in code reviews.

The kernel also gives you centralized timing. Instead of maintaining a tick counter and comparing deltas in your super-loop, you use vTaskDelayUntil() for precise periodic execution, software timers for callbacks, and xTaskNotify for lightweight signaling. If you plan to evaluate kernels, my comparison in RTOS Fundamentals: FreeRTOS vs Zephyr for Embedded Projects walks through the practical differences in scheduling, memory models, and driver ecosystems between the two most common choices.

Latency, Determinism and the Real Cost of Context Switching

This is where intuition often fails. An RTOS does not improve worst-case interrupt latency — it adds to it. On a bare metal system, the latency from an interrupt firing to the first instruction of your ISR is just hardware stacking plus your vector table jump: often 12-20 cycles on Cortex-M4. With an RTOS, you may also have kernel critical sections where interrupts are masked or deferred, and if the ISR calls a kernel API like xQueueSendFromISR(), you may pend a context switch that will not run until the kernel exits its critical section.

For most IoT sensing applications, that extra 1-5 microseconds does not matter. Where it does matter is closed-loop control. In one motor controller project, I measured an additional 2.8 microseconds of jitter in the current-loop ISR when the RTOS tick was set to 1 kHz with frequent queue operations. We solved it by moving the current loop to a highest-priority interrupt above configMAX_SYSCALL_INTERRUPT_PRIORITY so it could never be masked by the kernel, and keeping it entirely outside the RTOS API — an approach explicitly recommended for hard real-time ISRs in the Zephyr and FreeRTOS porting guides.

Measuring What Actually Matters

Do not rely on marketing numbers. Measure on your MCU with your compiler. I use a GPIO toggle at ISR entry/exit and a logic analyzer to capture:

1. Interrupt latency: Time from external pin edge to ISR GPIO high.
2. Context switch time: Time from an ISR signaling a high-priority task to that task's first instruction (toggle another pin at task entry). On an 80MHz STM32L4 with FreeRTOS, I typically see 1.5-3 microseconds for a switch if no FPU state needs saving, and 4-7 microseconds with FPU.
3. Tick jitter: Variation in periodic task wake-up when the system is loaded.

If your application needs sub-50 microsecond deterministic response, keep that path in a bare metal ISR or DMA-driven and let the RTOS handle everything else. For time-stamped sensor sampling at 100Hz or MQTT publishing at 1Hz, the scheduler's overhead is negligible. The trade-off is not raw speed, but predictability under load — bare metal is predictable because you control everything manually; RTOS is predictable because you assign priorities and let the kernel enforce them, provided you configure priorities and stack sizes correctly.

Memory Footprint and Toolchain Complexity: Where Small MCUs Force Your Hand

Memory is often the deciding factor. A minimal bare metal project on GCC with newlib-nano might occupy 4-8KB Flash and a few hundred bytes RAM beyond your own buffers. Adding FreeRTOS with one or two tasks costs roughly 4-6KB Flash for the kernel plus 1-2KB RAM for its heap and TCBs, and each task needs its own stack — typically 256-512 bytes minimum for a small task, 1KB+ for anything using printf or TLS. Zephyr, with its integrated drivers and device tree, is larger still: a minimal Zephyr blinky is ~30-50KB Flash, but you get a full driver framework and networking stack in return. The Zephyr Project Documentation provides detailed memory footprint tables by board that are worth checking before you commit.

RAM is the tighter constraint. I would not consider an RTOS on a part with less than 16KB RAM unless the vendor provides a heavily trimmed configuration. On an STM32G0 with 8KB RAM, bare metal leaves room for buffers and a proper packet queue. On an nRF52840 with 256KB RAM or an ESP32-S3 with 512KB, the kernel's overhead is noise compared to your TLS and MQTT buffers. For help estimating static versus dynamic needs, see Memory Management in Embedded C: Static Allocation, Pool and Arena Patterns — I apply those same pool patterns to size RTOS heaps and avoid fragmentation.

Toolchain and Debug Overhead

Bare metal debugging with SWD is simple: breakpoints, single-step, view registers. An RTOS adds state you need visibility into. Modern GDB + OpenOCD or J-Link can show task lists, stack high-water marks, and queue contents, but you need an RTOS-aware plugin. I rely on uxTaskGetStackHighWaterMark() during development and HardFault handlers that dump the offending task's stack. Without that, a stack overflow in one task silently corrupts another and you will spend days chasing ghosts. If you have not set up RTOS-aware debugging, the guide on Embedded Debugging with JTAG and SWD: Tools, Techniques and Workflows covers the probe and IDE configuration I use daily.

// Detect stack overflow early during development
// FreeRTOSConfig.h
#define configCHECK_FOR_STACK_OVERFLOW 2
#define configUSE_MALLOC_FAILED_HOOK   1

void vApplicationStackOverflowHook(TaskHandle_t xTask, char *pcTaskName) {
    // Catch overflow immediately: log task name and halt
    printf("STACK OVERFLOW in %s\r\n", pcTaskName);
    taskDISABLE_INTERRUPTS();
    for(;;);
}

void vApplicationMallocFailedHook(void) {
    printf("HEAP ALLOCATION FAILED\r\n");
    taskDISABLE_INTERRUPTS();
    for(;;);
}

// In a debug CLI task: periodically report usage
void print_task_stats(void) {
    UBaseType_t watermark = uxTaskGetStackHighWaterMark(NULL);
    printf("Stack free: %u words\r\n", (unsigned)watermark);
}

Build complexity also rises. With bare metal you compile your files and link. With FreeRTOS you add the kernel sources and a FreeRTOSConfig.h; with Zephyr you adopt west, device tree, and Kconfig. That upfront investment pays back quickly on large projects through driver reuse, but for a 2KB utility firmware that just blinks LEDs and reads a pin, it is pure overhead.

Deciding Between Bare Metal and RTOS for Connected IoT Products

For IoT, connectivity changes the calculus. Once you add a radio, you inherit asynchronous, long-running operations: joining a network, handling TLS handshakes that take seconds, retrying MQTT publishes, and parsing AT command responses from a modem. Trying to multiplex those in a super-loop leads to blocking delays or complex fragmented state machines. An RTOS lets you isolate each subsystem.

Here is how I decide in practice:

Choose bare metal when: you have a single real-time job, deep sleep with sub-10uA targets, less than ~32KB Flash / 8KB RAM, or hard real-time control where you need cycle-accurate guarantees and can prove timing by inspection. Examples include battery-powered sensor beacons, simple actuators, and motor drivers.

Choose an RTOS when: you have three or more concurrent activities (especially mixing periodic sampling, user interface, and networking), need blocking communications stacks (TCP/IP, MQTT, BLE GATT), want to reuse vendor middleware written for RTOS, or have a team where different developers own different subsystems. Examples include gateways, wearables with displays, and any device using MQTT Specification over TLS where publish, subscribe, and heartbeat must run independently without blocking sampling.

Criterion Bare Metal (Super-Loop + ISRs) RTOS (FreeRTOS / Zephyr)
Typical RAM Budget < 8–16 KB viable; hundreds of bytes overhead 16 KB minimum practical; 4–8 KB kernel + stacks
Typical Flash Overhead 2–8 KB 6–50 KB depending on kernel & drivers
Concurrency Model Manual flags, state machines, non-blocking checks Tasks, queues, mutexes, event groups
Worst-Case ISR Latency Lowest possible (hardware only) Slightly higher due to critical sections
Timing Precision Manual; drift if loop timing slips vTaskDelayUntil / kernel timers; priority-driven
Power Management Explicit WFI / stop modes; full control Tickless idle supported; requires port integration
Debuggability Simple stack traces; logic analyzer friendly Needs RTOS-aware GDB; stack overflow risks
Best For Simple sensors, motor control, ultra-low-power Gateways, connected products, multi-protocol

One nuance: using an RTOS does not automatically give you low power. You must enable tickless idle (configUSE_TICKLESS_IDLE in FreeRTOS or CONFIG_PM in Zephyr) so the kernel can stop the SysTick and enter STOP mode when all tasks are blocked. Without it, the 1kHz tick will wake the core every millisecond and ruin your battery life. In my last nRF52 project, enabling tickless idle dropped average sleep current from 1.2mA to 18uA. Bare metal had a slight edge because I could sleep with zero periodic wake-ups, but the difference was small enough to justify the RTOS for networking reasons.

Migrating Paths: Starting Bare Metal and Growing Into an RTOS Without a Rewrite

If you are unsure, start bare metal but structure the code so a later move to RTOS does not require a rewrite. I organize all hardware access behind thin driver interfaces that do not assume a concurrency model. For example, an I2C driver exposes i2c_transfer_nb() that starts a transfer and returns immediately, and a completion callback or flag. The super-loop polls completion; an RTOS task would block on a semaphore given from the ISR. The driver itself does not need to change.

Second, never call blocking delays inside drivers. Use explicit timeout parameters and return error codes. Third, keep ISRs minimal: just move data to a ring buffer and signal. Those three habits make porting straightforward. When I migrated a smart irrigation controller from bare metal to FreeRTOS mid-project because we added Wi-Fi provisioning, the driver layer stayed identical — I simply wrapped the super-loop state machines into tasks and replaced flag polling with xQueueReceive and xEventGroupWaitBits.

The opposite direction — stripping an RTOS out — is much harder, because RTOS code tends to spread vTaskDelay and queue dependencies throughout business logic. If you anticipate that cost or certification may later force you back to bare metal, avoid deep coupling to kernel APIs in your application logic by isolating RTOS calls in a small porting layer.

In the end, the question is not which is more advanced, but which lets you meet deadlines, hit power budgets, and keep the firmware maintainable for the people who will support it next year. For simple, high-volume sensors I still reach for bare metal without hesitation. For anything that talks to the cloud and must do two things at once, an RTOS pays for itself within weeks. Build a prototype both ways on your target MCU if you have time — measure current, latency, and RAM — and the numbers will make the decision for you.

Frequently Asked Questions

Can I use an RTOS on a low-end MCU like an ATtiny or STM32F0 with 16KB Flash?

You can, but you probably should not. With 16KB Flash and 4-8KB RAM, the kernel and task stacks leave very little room for application code and buffers. In that range I stay with bare metal and use SysTick-driven state machines. If you need richer concurrency, step up to an STM32G4 or nRF52 where the overhead is negligible. If you must use an RTOS on a small part, choose a minimal configuration like FreeRTOS with configMINIMAL_STACK_SIZE tuned down and only one or two tasks.

Does an RTOS guarantee real-time behavior?

No kernel by itself guarantees it — your design does. An RTOS is more accurately described as a real-time kernel: it provides deterministic scheduling primitives, but you must assign priorities correctly, bound critical sections, avoid priority inversion with mutexes, and ensure high-priority tasks have sufficient stack and CPU budget. Rate-monotonic analysis and measuring worst-case execution time still matter. For hard deadlines below a few tens of microseconds, keep the critical path in a top-priority ISR outside the kernel.

Is bare metal faster than RTOS?

For raw interrupt latency and throughput in a tight loop, bare metal is marginally faster because there is no context switch or tick interrupt. In practical IoT workloads that spend most time waiting for I/O, the difference is dominated by peripheral and network time, not CPU cycles. I have seen RTOS-based designs outperform bare metal equivalents because the blocking model allowed better use of DMA and sleep while waiting, rather than polling.

How do I safely share data between interrupts and tasks, or between tasks?

Never share mutable data without protection. Between an ISR and a task, use lock-free ring buffers or queue APIs designed to be called from ISRs (like xQueueSendFromISR). Between tasks, use mutexes for shared peripherals (I2C, SPI) and queues or message buffers for data passing. Avoid disabling interrupts globally as a general lock — it increases latency. And always size queues for worst-case bursts, not average throughput, so you do not drop samples under load.

Related Articles

References & Standards: FreeRTOS Documentation · Zephyr Project Documentation · MQTT Specification