When I shipped my first Rust firmware to a fleet of Cortex-M4 sensor nodes after a decade of writing embedded C, the most surprising thing wasn't the language syntax — it was the silence. No mysterious hard faults after three weeks in the field, no corrupted sensor buffers that only appeared when the radio interrupt fired at the wrong moment. That silence is what memory safety sounds like on a microcontroller. Embedded Rust, built around the no_std ecosystem, gives you deterministic control over hardware without giving up the compiler's guarantees against entire classes of bugs that have plagued embedded C for decades. In this article, I'll walk through how I approach embedded Rust in production, from toolchain setup to concurrency and incremental adoption, with the practical tradeoffs I've learned on real hardware.
Why C's Memory Model Breaks Down on Long-Lived Microcontroller Deployments
In my experience, the bugs that cost the most in IoT deployments aren't the ones you catch on the bench. They are use-after-free errors in a packet queue, buffer overflows in a vendor driver, or data races between a UART ISR and the main loop that corrupt a shared struct once every 10,000 packets. On a desktop OS, the process crashes and restarts. On a battery-powered sensor running for years on a rooftop, it becomes a silent failure that drains the battery or reports corrupt data.
Rust's ownership and borrowing rules address this at compile time. Every value has a single owner, and the compiler tracks how references are borrowed. Mutable aliasing — having two mutable references to the same memory — is simply not allowed to compile. For microcontroller work, this is profound. Most of the embedded C patterns I relied on for safety, like carefully documented ownership of DMA buffers or disabling interrupts around shared accesses, become enforceable by the compiler instead of conventions in comments.
This doesn't mean Rust prevents logic errors. You can still write an incorrect PID loop or misconfigure a clock tree. What it does prevent is undefined behavior from memory violations. Null pointer dereferences, double frees, and out-of-bounds accesses are caught before the binary is ever flashed. For a deeper look at how these problems are traditionally managed in C, the techniques in Memory Management in Embedded C: Static Allocation, Pool and Arena Patterns remain relevant — Rust automates many of those same patterns with zero runtime cost.
The True Cost of Undefined Behavior at the Edge
I've found that teams often underestimate the cost of undefined behavior in field devices. The C standard leaves behavior undefined to allow optimization, but on an ARM Cortex-M, undefined behavior often manifests as a HardFault with a stacked program counter that points nowhere useful. Debugging that without a full trace, especially on devices where you cannot easily attach a debugger as described in Embedded Debugging with JTAG and SWD: Tools, Techniques and Workflows, can take days. Rust's approach is to make the invalid state unrepresentable. For example, peripheral access crates (PACs) generated from SVD files expose registers as typed structures where you cannot write an invalid bit pattern without using an explicit unsafe block. The compiler forces you to acknowledge the risk.
Taming the no_std Toolchain: Targets, Runtimes and Linker Configuration for Cortex-M
Standard Rust assumes an operating system, a heap allocator, and a standard library. A microcontroller has none of those by default. The no_std attribute tells the compiler not to link the standard library, and you build for a bare-metal target triple like thumbv7em-none-eabihf for Cortex-M4F or thumbv6m-none-eabi for Cortex-M0+.
Your project needs three pieces beyond your application code: a runtime crate, a linker script, and a target configuration. I use cortex-m-rt for the startup code and vector table, and cortex-m for core peripherals and assembly intrinsics. The memory layout is still defined by a linker script, just like in C.
Target Triples, Linker Scripts and Memory.x
Here's the minimal configuration I've settled on for an STM32L4 project. It is explicit and checked into version control, which I strongly recommend over relying on implicit defaults.
# Cargo.toml
[package]
name = "sensor-node"
version = "0.1.0"
edition = "2021"
[dependencies]
cortex-m = "0.7"
cortex-m-rt = "0.7"
cortex-m-rtic = "1.1" # optional, for scheduling
panic-probe = { version = "0.3", features = ["print-defmt"] }
defmt-rtt = "0.4"
[profile.release]
lto = true
opt-level = "s"
codegen-units = 1
panic = "abort"
# .cargo/config.toml
[build]
target = "thumbv7em-none-eabihf"
[target.thumbv7em-none-eabihf]
runner = "probe-rs run --chip STM32L475VG"
The actual RAM and flash regions are defined in a memory.x file that the linker includes. This is where you prevent the linker from placing data in non-existent memory, a mistake that is frustratingly common when porting between variants in the same family.
/* memory.x for STM32L475VG - 512K Flash, 128K RAM */
MEMORY
{
FLASH : ORIGIN = 0x08000000, LENGTH = 512K
RAM : ORIGIN = 0x20000000, LENGTH = 128K
}
/* Stack top - placed at end of RAM */
_stack_start = ORIGIN(RAM) + LENGTH(RAM);
With this setup, cargo build --release produces a correct ELF with the vector table at 0x08000000. I've found that keeping memory.x per-board and using Cargo features to select the right PAC (e.g., stm32l4 = { version = "0.15", features = ["stm32l475"] }) avoids a lot of configuration drift as the hardware evolves.
Ownership in 64KB: Making the Borrow Checker Work for SRAM-Constrained Firmware
The biggest mental shift when moving to embedded Rust is learning to work with the borrow checker without a heap. On a device with 64KB of SRAM, you rarely want to use dynamic allocation anyway. Rust excels here because ownership works perfectly with static allocation.
Instead of malloc, you create buffers as static objects or stack-allocated arrays and move ownership of them. For example, a driver that needs a transmit buffer doesn't take a pointer and a length; it takes ownership of a &mut [u8] and returns it when done. The compiler ensures only one part of the code can mutate that buffer at a time.
Where you do need shared, mutable state — and you will — Rust gives you explicit, auditable tools. For global state accessed from both main and interrupts, I use the patterns provided by cortex-m rather than raw static mut, which is unsafe and should be avoided.
Zero-Cost Abstractions vs. Hidden Allocations
A common concern is that Rust abstractions will bloat the binary or add hidden costs. In my measurements on Cortex-M4, this is largely unfounded if you are disciplined. Iterators, enums with pattern matching, and Result-based error handling all compile down to the same assembly as hand-written C loops and error codes when optimizations are enabled. What you must watch for is implicit use of alloc.
Crates marked no_std will not use the heap, but some crates that support both std and no_std will pull in alloc if you enable the alloc feature. On a heapless system, that will fail to link unless you provide an allocator, which is often not what you want. I audit dependencies with cargo tree and prefer crates from the heapless ecosystem — heapless::Vec, heapless::String, and bbqueue — which provide fixed-capacity collections backed by arrays. They make the memory cost visible at compile time, much like the static pool patterns used in robust C firmware.
Sharing Peripherals Safely: How Rust Enforces Correct Interrupt and DMA Access
Concurrency on a microcontroller is fundamentally about interrupts and DMA. In C, sharing a peripheral register block or a buffer between main context and an ISR is error-prone. You disable interrupts, access the data, and re-enable them, hoping you covered every path. Rust's type system encodes this directly.
The canonical pattern is the peripheral singleton. The PAC provides a Peripherals::take() method that can only be called once and returns ownership of all peripherals. You then split that struct and move each peripheral to the driver that owns it. The compiler prevents two drivers from taking ownership of the same UART.
use cortex_m::interrupt::Mutex;
use cortex_m::peripheral::Peripherals as CorePeripherals;
use stm32l4::Peripherals as DevicePeripherals;
use core::cell::RefCell;
static SHARED_COUNTER: Mutex<RefCell<u32>> = Mutex::new(RefCell::new(0));
#[entry]
fn main() -> ! {
let _core = CorePeripherals::take().unwrap();
let device = DevicePeripherals::take().unwrap();
// Move ownership of USART2 to the serial driver
let _serial = device.USART2;
loop {
cortex_m::interrupt::free(|cs| {
let mut counter = SHARED_COUNTER.borrow(cs).borrow_mut();
*counter += 1;
});
}
}
#[interrupt]
fn USART2() {
cortex_m::interrupt::free(|cs| {
let mut counter = SHARED_COUNTER.borrow(cs).borrow_mut();
// Safe, atomic access - interrupt is already at higher priority
// but the Mutex still enforces correct borrowing
*counter += 10;
});
}
In this example, Mutex<RefCell<T>> combined with cortex_m::interrupt::free provides a critical section that disables interrupts for the duration of the borrow. The borrow checker ensures you cannot forget to enter the critical section — you cannot access the data without the cs token that proves interrupts are disabled. For more complex scheduling, frameworks like RTIC (Real-Time Interrupt-driven Concurrency) build on this model to provide compile-time deadlock and priority inversion checks, which is a significant step beyond what you get from traditional RTOS primitives. If you are evaluating scheduling options, the tradeoffs discussed in RTOS Fundamentals: FreeRTOS vs Zephyr for Embedded Projects are a useful baseline, and RTIC can be seen as a Rust-native alternative for interrupt-driven designs. For low-level ISR design principles that apply regardless of language, I still refer to Interrupt Handling Best Practices: Priority, Latency and ISR Design.
The Zephyr Project Documentation provides a good reference for how C-based systems handle similar concurrency through kernel objects and interrupt locking — Rust's model achieves the same isolation but moves the verification to compile time.
Incremental Adoption: Running Rust and C Together in Production Firmware
Very few teams can rewrite a 100,000-line C codebase overnight. In my experience, the most successful Rust migrations are incremental. Rust has excellent FFI support, and you can call C from Rust and vice versa with minimal overhead.
The pattern I use is to keep the existing C startup, FreeRTOS kernel, or vendor HAL in place and add new modules in Rust. You define extern "C" functions in Rust that are callable from C, and you declare external C functions in Rust to call them. The key is to be strict about the boundary.
At the FFI boundary, you must use unsafe. This is correct — you are telling the compiler "I have verified that this C code upholds Rust's invariants." Keep these blocks small and wrap them in safe abstractions as quickly as possible. For example, if you have a proven C driver for a complex radio, wrap its init, send, and isr_handler functions in a safe Rust struct that owns the underlying C state pointer and exposes a safe method.
I've found that starting with peripheral drivers or protocol parsers is ideal. Parsers are particularly good candidates because Rust's nom or defmt parsers are memory-safe by construction, eliminating a major source of vulnerabilities in C-based MQTT or CBOR parsers. The MQTT Specification is a good example where a Rust parser can guarantee it never reads beyond the packet buffer, even on malformed input, without adding runtime checks beyond what you'd write in C anyway.
For build integration, I compile the Rust code as a static library with crate-type = ["staticlib"] and link it into the existing CMake or Make-based C build. Tools like cbindgen generate the C headers automatically from Rust types, keeping the interface in sync.
Panic Handling and Footprint Control When Every Kilobyte Counts
On a desktop, a panic can unwind the stack and print a backtrace. On a Cortex-M with 128KB of flash, you need to decide explicitly what happens. Rust gives you two panic strategies: abort and unwind. For embedded targets, you almost always want abort, which immediately halts or resets. Unwinding requires extra metadata and is not supported on thumbv6m anyway.
Even with abort, you must provide a panic handler. In development, I use panic-probe with defmt to log panics over RTT — it gives you file, line, and message at near-zero cost. For production, I switch to panic-halt or a custom handler that logs the panic location to backup SRAM and triggers a watchdog reset. Never let a panic handler return; it must be -> ! (diverging).
Binary size is often cited as a concern, but modern Rust is competitive with C when configured correctly. My release profile disables debug assertions, enables LTO, and optimizes for size. The table below summarizes what I've measured on a recent STM32L4 sensor application comparing equivalent functionality.
| Configuration | Flash Usage (approx.) | Runtime Overhead | Debug Experience |
|---|---|---|---|
| C with -Os, no LTO | 42 KB | None | Limited HardFault info |
| Rust panic-halt, opt-level="s", LTO | 44 KB | None (abort on panic) | Minimal (halt only) |
| Rust panic-probe + defmt-rtt | 48 KB | ~200 bytes RAM for RTT buffer | File/line + formatted logs |
| Rust with heapless + RTIC | 51 KB | Deterministic scheduling, no heap | Full RTIC task tracing |
As you can see, the delta is 2-6KB for the core benefits. The largest increase comes from debug infrastructure that you would remove or minimize in a production build. I've found that the flash cost is quickly recovered by eliminating defensive code and manual error checks that C requires. The FreeRTOS Documentation shows similar footprint numbers for its kernel — Rust's core runtime is in the same range.
To keep size down, avoid format! in hot paths, prefer defmt over core::fmt, and use the cargo-bloat tool to find unexpected dependencies pulling in large code. One large driver crate can easily add 20KB if you enable all its default features.
Frequently Asked Questions
Can I use Embedded Rust without a heap allocator?
Yes, and most production firmware does. The entire no_std ecosystem is designed to work without alloc. You use stack-allocated arrays, static buffers, and crates like heapless that provide fixed-capacity data structures. A heap allocator like linked_list_allocator is only needed if you explicitly want dynamic allocation, which I avoid on devices with less than 64KB RAM because it introduces fragmentation risk.
How does Rust interact with existing FreeRTOS or Zephyr tasks?
Rust can run on top of both. You can either write a Rust task that is spawned by the RTOS using its C API via FFI, or use a Rust-native executor like RTIC or Embassy. Embassy, for example, provides async executors that work cooperatively with interrupts and can even integrate with the Zephyr kernel. I've had good results wrapping FreeRTOS queues and semaphores in safe Rust abstractions and calling them from Rust tasks compiled as part of a larger C firmware.
Is unsafe Rust just as dangerous as C?
No, but it requires the same discipline. unsafe in Rust doesn't turn off the borrow checker; it only allows five specific operations like dereferencing raw pointers and calling unsafe functions. The crucial difference is that unsafe is scoped and explicit. You must mark the block, and safe code cannot cause undefined behavior regardless of what unsafe code does, provided the unsafe code correctly upholds invariants. In audits, I can search for every unsafe block; in C, every line is implicitly unsafe.
What is the best way to debug panic and HardFault in the field?
In development, use panic-probe and defmt-rtt with probe-rs to get live panic messages. For field devices, implement a custom panic handler that writes the panic location and a small stack dump to a no-init RAM section or external flash before resetting. Pair this with the Cortex-M's HardFault handler to capture the stacked registers. On reboot, send that buffer with your next telemetry packet. This has saved me weeks of debugging devices that were failing only at customer sites.