The shift from cloud-based ML inference to on-device processing is driven by three constraints that no amount of network optimization can solve: latency requirements under 10ms for real-time control loops, privacy regulations that prohibit transmitting raw sensor data, and connectivity limitations in LPWAN-connected deployments where uplink bandwidth is measured in bytes per second. Edge AI addresses all three by moving the model to where the data lives.
But running inference on a Cortex-M microcontroller with 256 KB of flash and 64 KB of RAM is a fundamentally different engineering challenge than serving a model from a GPU cluster. This article covers the practical techniques for getting useful ML inference running on embedded hardware without destroying your power budget or accuracy targets.
The TinyML Landscape
TinyML occupies the extreme end of the edge AI spectrum — inference on microcontrollers running at milliwatt power levels. The defining constraint is not compute speed but memory. A typical Cortex-M4 MCU has 256 KB of flash for the entire application (model weights, inference engine, application logic, and RTOS) and 64 KB of SRAM for runtime buffers.
Within these constraints, TinyML models handle specific tasks remarkably well:
- Keyword spotting: 30-50 KB models detecting wake words with 95%+ accuracy at under 1mW
- Anomaly detection: Autoencoder models under 20 KB identifying abnormal vibration patterns in industrial motors
- Gesture recognition: IMU-based classifiers under 15 KB for wearable applications
- Image classification: MobileNet variants quantized to 250 KB for visual inspection on ESP32-S3 with PSRAM
The key insight is that TinyML does not attempt to shrink general-purpose models. Instead, it uses task-specific architectures designed from scratch for the target hardware, trained on domain-specific datasets, and quantized with hardware-aware optimization.
Model Quantization: From FP32 to INT8 and Beyond
Quantization converts model weights and activations from 32-bit floating point to lower-precision representations. This reduces model size, speeds up inference (integer arithmetic is faster on MCUs), and lowers power consumption. The three main approaches offer different tradeoffs:
| Method | Precision | Size Reduction | Accuracy Impact | Workflow |
|---|---|---|---|---|
| Post-training dynamic | INT8 weights, FP32 activations | ~2x | Minimal | No retraining required |
| Post-training static | INT8 weights + activations | ~4x | 0.5-2% loss | Requires calibration dataset |
| Quantization-aware training | INT8 full pipeline | ~4x | <0.5% loss | Fine-tuning with fake quantization nodes |
| INT4/binary | 4-bit or 1-bit weights | ~8-32x | 2-10% loss | Specialized architectures required |
For production embedded deployments, quantization-aware training (QAT) is almost always worth the engineering investment. The process inserts fake quantization nodes during training so the model learns to compensate for precision loss. On RTOS-based systems, the resulting INT8 model runs 2-4x faster than FP32 on Cortex-M cores with DSP extensions, using the SMLAD and SMMLA instructions for fused multiply-accumulate operations.
Mixed-Precision Strategies
Not all layers tolerate quantization equally. The first and last layers of a CNN are typically more sensitive to precision loss than intermediate layers. Mixed-precision quantization keeps sensitive layers at INT16 or FP32 while quantizing the rest to INT8, achieving near-FP32 accuracy with most of the INT8 performance benefit.
# TensorFlow Lite converter with per-layer quantization
converter = tf.lite.TFLiteConverter.from_saved_model(model_path)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = calibration_generator
# Keep first conv layer at FP16
def quantize_annotate(layer):
if layer.name == 'conv2d_0':
return quantize.quantize_annotate_layer(layer,
quantize_config=Float16QuantizeConfig())
return layer
Hardware Accelerators: NPU, DSP, and Embedded GPU
Raw CPU inference on a Cortex-M4 at 168 MHz processes roughly 50 million multiply-accumulate operations per second (50 MMAC/s). A dedicated neural processing unit on the same die can achieve 1-10 GMAC/s — a 20-200x improvement for neural network workloads. Understanding when each accelerator type makes sense is critical for hardware selection.
NPU (Neural Processing Unit)
NPUs contain fixed-function hardware optimized for the operations that dominate neural network inference: 2D convolutions, matrix multiplications, and activation functions. The Arm Ethos-U55 is representative of the class, delivering up to 128 GMAC/s at under 50mW for INT8 workloads. NPUs excel when the model architecture aligns with the hardware's supported operations. Custom layers or unusual activation functions fall back to CPU execution, sometimes negating the NPU's throughput advantage.
DSP (Digital Signal Processor)
DSPs offer more programmability than NPUs at the cost of lower peak throughput per watt. For applications that combine signal processing (FFT, filtering, feature extraction) with ML inference, a DSP handles the entire pipeline without data movement between cores. The Cadence Tensilica HiFi and CEVA-BX series support neural network operations alongside traditional DSP workloads.
Embedded GPU
Embedded GPUs like the Mali-G series provide the highest raw compute but at significantly higher power. They shine in computer vision applications where parallel pixel processing aligns with GPU architecture. For battery-powered devices, the power overhead makes embedded GPUs impractical below the 500mW range.
Inference Frameworks: TensorFlow Lite Micro vs ONNX Runtime
Two frameworks dominate embedded inference. TensorFlow Lite Micro (TFLM) targets the extreme low end — bare-metal deployment on Cortex-M with no dynamic memory allocation. The interpreter reads a flatbuffer model and executes operations using pre-allocated tensor arenas. TFLM's footprint starts at approximately 20 KB of flash, making it viable on MCUs with as little as 64 KB.
ONNX Runtime Mobile targets slightly larger devices (Cortex-A class) and provides broader operator coverage and optimization passes like graph fusion and constant folding. Its minimum footprint is approximately 200 KB, requiring devices with at least 512 KB of flash and 128 KB of RAM.
| Feature | TF Lite Micro | ONNX Runtime Mobile |
|---|---|---|
| Target hardware | Cortex-M0+ to M7 | Cortex-A, RISC-V application |
| Minimum flash | ~20 KB | ~200 KB |
| Dynamic allocation | None (static arena) | Optional allocator |
| OS requirement | Bare metal / RTOS | RTOS / Linux |
| Model format | .tflite (flatbuffer) | .ort (optimized ONNX) |
| Quantization support | INT8, INT16 | INT8, FP16, INT4 |
| Hardware delegation | Via custom ops | Execution providers |
For MQTT-connected sensor nodes running keyword detection or anomaly classification, TFLM on a Cortex-M4 is the right choice. For smart camera applications running object detection on a Cortex-A53, ONNX Runtime provides better optimization and broader model support.
Power-Constrained Inference Patterns
On energy-harvesting devices, every microjoule counts. Three architectural patterns minimize inference energy consumption:
Cascaded Inference
A lightweight classifier (under 5 KB) runs continuously at minimal power, triggering a larger, more accurate model only when the initial classifier detects a potential event. For audio wake-word detection, the cascade reduces average power by 10-50x because the full model runs less than 1% of the time.
Duty-Cycled Processing
Instead of continuous inference, the device samples sensors at intervals and batches inference. A temperature monitoring system might sample every 30 seconds and run anomaly detection on a buffer of 10 readings, amortizing the wake-up and inference cost across multiple data points.
Early Exit Networks
Models with multiple exit points allow inference to terminate early when confidence exceeds a threshold. The first few layers handle easy cases (90% of inputs in many applications), and the full network processes only ambiguous inputs. This pattern reduces average inference energy by 30-60% in deployment.
Model Architecture Selection
Architecture choice determines whether a model fits on the target hardware. MobileNetV2 with depthwise separable convolutions remains the workhorse for embedded vision, offering configurable width and resolution multipliers to trade accuracy for size. EfficientNet-Lite variants achieve better accuracy per parameter but require more memory for activations. For time-series classification, 1D temporal convolutions and GRU cells outperform fully connected networks in both accuracy and efficiency.
When selecting architectures for production deployment, profile the model on the actual target hardware. Theoretical FLOP counts correlate poorly with actual inference time because memory bandwidth, cache behavior, and instruction pipeline stalls dominate performance on constrained devices.
Deployment and Validation
Deploying an edge AI model requires validation beyond accuracy metrics. Key considerations include over-the-air update mechanisms for model updates, deterministic inference timing for safety-critical applications, and graceful degradation when input data falls outside the training distribution.
For fleet monitoring, track inference latency distributions, confidence score histograms, and anomaly rates across firmware versions. A sudden shift in confidence distribution often indicates data drift before accuracy degrades — enabling proactive model retraining rather than reactive incident response.
Frequently Asked Questions
What is TinyML and how does it differ from traditional edge AI?
TinyML refers to machine learning inference running on microcontrollers with less than 1 MB of RAM and milliwatt-level power budgets. Traditional edge AI runs on application processors or GPUs with megabytes to gigabytes of memory. TinyML models are typically under 500 KB, use integer-only arithmetic, and target always-on sensing applications like keyword spotting or anomaly detection where battery life is measured in years.
How much accuracy loss should I expect from INT8 quantization?
Post-training INT8 quantization typically reduces model accuracy by 0.5-2% for classification tasks and 1-3% for detection tasks compared to FP32 baselines. Quantization-aware training narrows this gap to under 0.5% for most architectures. The accuracy impact depends on the model architecture — depthwise separable convolutions and attention mechanisms tend to be more sensitive than standard convolutions.
Which hardware accelerator should I choose for embedded AI?
NPUs provide the best performance-per-watt for fixed neural network operations. DSPs offer more flexibility for multi-stage signal processing pipelines. Embedded GPUs provide the highest raw throughput but consume more power. For battery-powered devices under 100mW, NPUs are typically the best choice.
Can I run transformer models on microcontrollers?
Small transformer models can run on Cortex-M7 MCUs with significant constraints. Models with 2-4 attention heads and embedding dimensions of 64-128 fit in 512 KB-1 MB of flash. Inference latency ranges from 50-500ms. For real-time applications, quantized CNN or RNN architectures remain more practical on sub-MHz class MCUs.
What is the minimum hardware needed to run TensorFlow Lite Micro?
TFLM requires approximately 20 KB of flash for the core interpreter and 4-16 KB of RAM for the tensor arena. An ARM Cortex-M3 at 48 MHz with 64 KB flash and 16 KB RAM represents the practical minimum. The framework supports bare-metal deployment, though running under an RTOS like FreeRTOS or Zephyr simplifies memory management.