ONNX Runtime for Edge: Cross-Framework Model Deployment on ARM

Moving a trained model from a workstation to an ARM-based edge device should be straightforward, but anyone who has tried to ship PyTorch to a Raspberry Pi, TensorFlow to an NXP i.MX8, or a scikit-learn pipeline to an NVIDIA Jetson Orin Nano knows the friction is real. Different frameworks, different runtime dependencies, and wildly different hardware accelerators turn what looks like a simple file copy into a week of cross-compilation and debugging. In my experience deploying vision and time-series models to Cortex-A53, A72, and A78 fleets, ONNX Runtime has become the most practical abstraction for this problem. It decouples model authoring from model execution, gives you a single C++ API that runs on everything from a Yocto-based gateway to an ARM64 server, and lets you swap execution providers without rewriting application code. This article walks through how I use ONNX Runtime to standardize cross-framework deployment on ARM, where the pitfalls are, and how to get predictable latency and memory usage in production.

Why ONNX Runtime Solves the ARM Edge Deployment Fragmentation Problem

Most teams start with framework-native runtimes. You train in PyTorch, then try to run TorchScript on the device. Or you train in TensorFlow and ship TFLite. That works until you have three model teams using two frameworks targeting four hardware SKUs. Suddenly you are maintaining separate inference stacks, separate quantization toolchains, and separate CI pipelines for each combination. ONNX (Open Neural Network Exchange) addresses this by providing an intermediate graph representation. ONNX Runtime is the engine that executes that graph.

What makes it valuable on ARM is not just portability, but the execution provider architecture. The runtime itself is small and CPU-agnostic, but at session creation you bind it to an accelerator backend: default CPU, XNNPACK, ARM Compute Library (ACL), Arm NN, or even a vendor-specific NPU provider. Your application code does not change when you move from a Cortex-A53 gateway with no accelerator to a Rockchip RK3588 with an NPU. You only change the session options. I've found that this separation saves more time than any single optimization, because it lets the embedded team own the deployment configuration while the ML team continues to iterate in their preferred framework.

Where ONNX Fits in a Typical Edge Stack

In a production gateway I recently worked on, the stack was: Yocto Kirkstone on a quad-core Cortex-A72, Zephyr on a companion Cortex-M4 for sensor polling, and MQTT for cloud sync. The Linux side ran ONNX Runtime as a system service written in C++. Models were delivered as versioned .onnx files over MQTT, validated on disk, and loaded without restarting the service. The abstraction meant we could run a MobileNetV2 for image inspection and a temporal convolutional network for vibration analysis side-by-side in the same process, each with its own thread pool and provider. For teams using Zephyr Project Documentation for their microcontroller firmware, this split — Zephyr for deterministic I/O and Linux + ONNX Runtime for ML — has proven far more maintainable than trying to cram inference onto the MCU.

ONNX Runtime also handles the unglamorous details that break deployments: dynamic shapes, custom ops, and graph optimizations. When you enable ORT_ENABLE_ALL, it folds constants, eliminates dead nodes, and fuses operators like Conv+BatchNorm+ReLU before any device-specific acceleration runs. That alone cut our graph initialization time by 30% on cold boot.

Converting PyTorch and TensorFlow Models to ONNX for ARM Targets

The conversion step is where most projects lose accuracy or introduce silent failures. Exporting to ONNX is not just saving a file; you are tracing or scripting your model into a static graph with a specific opset version. For ARM edge deployment, I standardize on opset 17 or 18, as they have mature support in ONNX Runtime 1.16+ and cover operators like LayerNorm and GELU without needing custom ops.

For PyTorch, use torch.onnx.export with a real dummy input that matches your deployment resolution and data layout. Avoid exporting with batch size hard-coded. Define dynamic axes for batch, and verify the exported graph with onnx.checker and onnxruntime inference on your development machine before cross-compiling. For TensorFlow, the path is tf2onnx. I've had more issues here with control flow and string ops, so I keep TensorFlow models to pure functional APIs before conversion. If you use tf.keras with custom layers, write a small Python test that runs the same input through both the original framework and the ONNX graph and compares outputs within 1e-4 tolerance.

# Export PyTorch model to ONNX with dynamic batch axis for ARM deployment
import torch
import onnx

# Assume 'model' is your trained PyTorch model in eval mode
model.eval()
dummy_input = torch.randn(1, 3, 224, 224)

torch.onnx.export(
    model,
    dummy_input,
    "mobilenetv2_arm.onnx",
    export_params=True,
    opset_version=17,
    do_constant_folding=True,
    input_names=['input'],
    output_names=['output'],
    dynamic_axes={'input': {0: 'batch_size'}, 'output': {0: 'batch_size'}}
)

# Validate the graph
onnx_model = onnx.load("mobilenetv2_arm.onnx")
onnx.checker.check_model(onnx_model)
print("ONNX graph validated, opset:", onnx_model.opset_import[0].version)

Common Conversion Failures I See on ARM Projects

Three issues appear repeatedly. First, NCHW vs NHWC layout mismatches. PyTorch is NCHW, but many ARM accelerators prefer NHWC internally. ONNX Runtime handles layout conversion automatically for the CPU provider, but if you later enable ACL, a transpose that was free can become a layout bottleneck. Second, unsupported ops like NonZero or Unique that force a fallback to CPU and kill performance. Run onnxruntime.tools.check_onnx_model to list kernels not covered by your chosen execution provider. Third, large constants embedded as initializers bloat the file for OTA updates. I externalize weights with onnx.save_model(..., save_as_external_data=True) when the .onnx exceeds 40MB, which helps when you deliver models over constrained links as described in AWS IoT Greengrass: Edge Computing, Local Processing and Cloud Sync.

If your model includes pre-processing like normalization or resizing, decide whether that belongs inside the graph. For edge cameras, I bake normalization into the graph as Mul and Add nodes so the C++ code just feeds raw uint8 buffers. For vibration data, I keep scaling outside the graph because calibration constants change per device and are easier to update in application code.

Building ONNX Runtime from Source for Cortex-A and Cortex-M Adjacent Platforms

Prebuilt ONNX Runtime wheels for ARM64 (aarch64) work fine for Raspberry Pi OS and Ubuntu, but production gateways running Yocto, Buildroot, or custom Debian rarely have the luxury of using them. In my experience, building from source is worth the initial setup because you can trim the binary size by 60% and enable exactly the execution providers you need. A default build with all providers enabled is over 120MB. A minimal build with only CPU and XNNPACK is under 30MB, which matters when your rootfs is 256MB.

For a Cortex-A72 target, I cross-compile from an x86_64 build host using the ARM GCC 11 toolchain. The key CMake flags are --arm --config Release --build_shared_lib --parallel --use_xnnpack --skip_tests. Disable training, disable Python bindings if you only need C++, and set CMAKE_SYSTEM_PROCESSOR correctly. If you need ACL, add --use_acl and point to a prebuilt Compute Library. For 32-bit Cortex-A7/A53 boards that still run armv7l, add --arm --armhf and verify NEON is enabled — ONNX Runtime will fall back to slow scalar kernels without it.

// Minimal ONNX Runtime C++ inference on ARM64 Linux with XNNPACK
#include <onnxruntime_cxx_api.h>
#include <vector>
#include <iostream>

int main() {
    Ort::Env env(ORT_LOGGING_LEVEL_WARNING, "edge-inference");
    Ort::SessionOptions session_options;
    
    // Tune for ARM big.LITTLE: 2-4 threads is usually optimal
    session_options.SetIntraOpNumThreads(4);
    session_options.SetGraphOptimizationLevel(GraphOptimizationLevel::ORT_ENABLE_ALL);
    session_options.SetExecutionMode(ExecutionMode::ORT_SEQUENTIAL);

    // Prefer XNNPACK on ARM; falls back to CPU if unavailable
    std::vector<std::string> providers = Ort::GetAvailableProviders();
    try {
        session_options.AppendExecutionProvider("XNNPACK");
    } catch (const Ort::Exception& e) {
        std::cout << "XNNPACK not available, using default CPU: " << e.what() << std::endl;
    }

    Ort::Session session(env, "mobilenetv2_arm.onnx", session_options);

    // Query input shape
    Ort::AllocatorWithDefaultOptions allocator;
    auto input_name = session.GetInputNameAllocated(0, allocator);
    std::vector<int64_t> input_shape = {1, 3, 224, 224};
    
    // Allocate input tensor (example: float32)
    size_t input_tensor_size = 1 * 3 * 224 * 224;
    std::vector<float> input_tensor_values(input_tensor_size, 0.5f);
    auto memory_info = Ort::MemoryInfo::CreateCpu(OrtArenaAllocator, OrtMemTypeDefault);
    Ort::Value input_tensor = Ort::Value::CreateTensor<float>(
        memory_info, input_tensor_values.data(), input_tensor_size, input_shape.data(), 4);

    auto output_name = session.GetOutputNameAllocated(0, allocator);
    const char* input_names[] = {input_name.get()};
    const char* output_names[] = {output_name.get()};

    auto output_tensors = session.Run(Ort::RunOptions{nullptr}, input_names, &input_tensor, 1, output_names, 1);
    float* output_data = output_tensors.front().GetTensorMutableData<float>();
    std::cout << "Inference complete, output[0]=" << output_data[0] << std::endl;
    return 0;
}

Memory and Threading Tuning for Real Hardware

On ARM you cannot treat threading as "more is better." I've measured peak throughput on a quad-core A72 with intra_op_num_threads=4 actually being slower than 2 threads for MobileNet-class models due to cache thrashing and thermal throttling. Start with 2, profile with ORT profiling enabled (session_options.EnableProfiling("profile.json")), and measure sustained FPS over 10 minutes, not just the first 100 inferences. For memory, pre-allocate tensors and reuse the Ort::Value objects across inferences instead of allocating per frame. On a 1GB RAM device running a 224x224 vision model, this reduced heap fragmentation enough to avoid OOM kills after 48 hours. If you are working close to microcontroller limits, review TinyML: Running Neural Networks on Microcontrollers with TensorFlow Lite to decide whether ONNX Runtime (which needs ~32MB RAM minimum) is appropriate or whether you should delegate the smallest models to TFLite Micro on the companion MCU.

Accelerating Inference with ARM NN, ACL and XNNPACK Execution Providers

The choice of execution provider determines your latency more than model architecture does on ARM. In my benchmarks across i.MX8M Plus, RK3568, and Raspberry Pi 4, the differences were stark. XNNPACK is the most reliable accelerator for quantized and float models on Cortex-A; it is maintained by Google, tightly optimized for NEON, and integrates with ONNX Runtime without external dependencies. ACL (ARM Compute Library) gives lower latency for float32 convolutions on larger inputs but adds a 10-15MB library and longer build times. Arm NN is now in maintenance mode for many vendors, so I avoid it for new designs unless a BSP requires it.

Execution Provider Best For Supported Precision Typical Latency (MobileNetV2, 224x224, Cortex-A72 1.8GHz) Dependencies
CPU (default) Compatibility, debugging FP32 185 ms None
XNNPACK General purpose ARM acceleration, quantized models FP32, INT8 (QDQ) 68 ms Built-in, no external lib
ACL (Arm Compute Library) FP32 conv-heavy vision models FP32, FP16 54 ms libarm-compute ~12 MB
Arm NN Legacy NPU/GPU delegation FP32, INT8 61 ms ArmNN + ACL + protobuf
NPU Vendor (e.g., rknn, Ethos-U) Maximum throughput, low power INT8 18-25 ms Vendor SDK, custom EP

To select a provider at runtime, query availability and set a fallback chain. I ship devices with XNNPACK as primary and CPU as fallback. On platforms where I have a verified NPU delegate, I try the NPU first, then XNNPACK, then CPU. Always log which provider was actually used — session.GetProviderType() after creation — because a silent fallback to CPU is the most common reason for "it worked in the lab but is 3x slower in the field."

For models that fuse well, like those that have been through Edge AI Model Optimization: Pruning, Quantization and Knowledge Distillation, XNNPACK with INT8 QDQ (Quantize-DeQuantize) format gives the best balance. The runtime keeps compute in INT8 and dequantizes only at the end, which preserves accuracy while using the fast NEON int8 kernels. I avoid the older QOperator format for new deployments; QDQ has better provider coverage and maps cleanly to ACL as well.

Quantizing ONNX Graphs for Memory-Constrained Edge Devices

On ARM edge hardware, quantization is not optional. A float32 ResNet50 at 98MB will not fit comfortably alongside an application, logging, and OTA buffers on a 512MB device. Static INT8 quantization typically reduces model size by 4x and latency by 2-3x with less than 1% accuracy drop if calibrated correctly. ONNX Runtime's quantization tool is Python-based and runs on your build host, not the device.

The workflow I use is: export FP32 ONNX, validate accuracy on a host, run static quantization with a representative calibration dataset (500-1000 samples from the field, not just training data), and then validate the INT8 graph again. For vision, include images with the actual lighting and blur you see on the device. For vibration analysis, include data from multiple operating speeds. Using a calibration set that is too clean is the fastest way to ship a quantized model that fails on edge data.

# Static INT8 quantization with ONNX Runtime (QDQ format) for ARM
from onnxruntime.quantization import quantize_static, QuantType, QuantFormat, CalibrationDataReader
import numpy as np
import glob
from PIL import Image

class ImageCalibrationReader(CalibrationDataReader):
    def __init__(self, image_dir, input_name="input"):
        self.image_files = glob.glob(f"{image_dir}/*.jpg")[:800]
        self.input_name = input_name
        self.index = 0
        self.preprocess = lambda img: np.expand_dims(
            np.transpose(np.array(Image.open(img).resize((224,224))).astype(np.float32) / 255.0, (2,0,1)), axis=0)

    def get_next(self):
        if self.index >= len(self.image_files):
            return None
        # Calibration must use same preprocessing as inference
        data = self.preprocess(self.image_files[self.index])
        self.index += 1
        return {self.input_name: data}

quantize_static(
    model_input="mobilenetv2_arm.onnx",
    model_output="mobilenetv2_arm_int8_qdq.onnx",
    calibration_data_reader=ImageCalibrationReader("./calibration_images"),
    quant_format=QuantFormat.QDQ,
    activation_type=QuantType.QInt8,
    weight_type=QuantType.QInt8,
    per_channel=True,  # critical for depthwise conv accuracy
    extra_options={'DedicatedPerChannel': True}
)
print("Quantized model saved for XNNPACK/ACL deployment")

What to Measure After Quantization

Do not trust size and latency alone. I run three checks on the quantized graph before deploying. First, accuracy delta on a holdout set that was not used for calibration — if top-1 drops more than 1.5%, try per-channel quantization or exclude the first and last layers from quantization. Second, operator coverage: load the INT8 model with the XNNPACK provider and confirm no nodes fell back to CPU FP32, which you can see in the session profiling log. Third, memory footprint on target: use /usr/bin/time -v or valgrind --tool=massif to capture peak RSS during 1000 inferences. I've seen quantized models that were faster but used more RAM due to extra dequantize nodes, which caused issues on devices with 256MB.

For time-series models used in anomaly detection, quantization behaves differently than vision. The dynamic range is often larger and outliers matter. In those cases I sometimes keep the LSTM or transformer backbone in FP16 and quantize only the embedding and classifier heads. This hybrid approach is covered in more detail in Predictive Maintenance with IoT and ML: Vibration Analysis and Anomaly Detection, and it aligns well with ONNX Runtime's ability to mix precisions per node.

Orchestrating Model Updates and Telemetry with MQTT at the Edge

An edge model is only useful if you can update it safely and monitor its behavior. I treat ONNX files as versioned artifacts delivered over MQTT, not as part of the firmware image. The firmware contains the ONNX Runtime engine and application logic; models are payloads. This separation lets me ship a model update in minutes without a full OTA, which is critical when you have 500 gateways on cellular connections.

The pattern that has worked reliably is: publish the new model to an MQTT topic like devices/{id}/model/update with a manifest containing SHA256, opset version, and required runtime version, as defined in the MQTT Specification. The device downloads to a staging file, verifies the hash, loads it in a temporary Ort::Session to validate that the graph initializes, and only then swaps it into production via an atomic file rename. Keep the previous model on disk for rollback. I also publish inference telemetry — latency p50/p95, provider used, and input drift metrics — back over MQTT to a shadow topic so the cloud can detect when a quantized model's accuracy is degrading due to data shift.

Practical Reliability Tips

Three lessons from field failures. First, enforce opset compatibility. A model exported with opset 18 will fail to load on a device running ONNX Runtime 1.14. Include the minimum runtime version in your model registry. Second, handle graceful degradation. If a new model fails to load, log the Ort::Exception, republish the previous model's hash as current, and alert. Never leave the device without a valid session. Third, account for atomicity with power loss. Use fsync after writing the staged model and keep two copies of the manifest. I've recovered devices that lost power mid-download because the boot logic could fall back to the last known good model instead of booting into a missing-file state.

For sensor fusion pipelines where inference consumes IMU and GPS data, sequence your pipeline so that ONNX Runtime inference runs in a dedicated thread with a bounded queue. The sensor fusion algorithm should not block on inference, and inference should not block on sensor I/O. This isolation prevents a slow inference (for example, when the NPU throttles thermally) from causing IMU sample drops that corrupt your state estimate.

Frequently Asked Questions

Can ONNX Runtime run on Cortex-M microcontrollers?

No, not directly. ONNX Runtime targets 32-bit and 64-bit application processors with an OS (Linux, Windows). It needs at least ~30MB RAM and a filesystem. For Cortex-M4/M7/M33, use TensorFlow Lite for Microcontrollers or ONNX Runtime's sister project, onnxruntime-micro, which is experimental and supports only a tiny operator set. In my designs, I run ONNX Runtime on the Cortex-A companion and use the MCU purely for sensor acquisition.

How do I debug a model that runs fine on x86 but fails on ARM?

Enable verbose Ort logging and compare the operator placement. On x86 you may be using OpenVINO or CUDA providers that support ops which XNNPACK or ACL on ARM do not. Run onnxruntime.tools.symbolic_shape_infer and dump the graph with onnxruntime --dump_model equivalent profiling. The most common culprits are dynamic control flow, large Reduce ops, and INT64 tensors that some ARM kernels do not support. Convert INT64 to INT32 where possible before export.

Should I use dynamic or static quantization for ARM edge deployment?

Use static quantization (with calibration data) for best latency on ARM. Dynamic quantization quantizes weights ahead of time but quantizes activations at runtime, which adds overhead and is less optimized in XNNPACK. Static QDQ INT8 gives you the fastest NEON kernels and deterministic latency. Only use dynamic quantization if you cannot collect a representative calibration set, and expect 10-20% higher latency than static.

What is the best way to validate accuracy after converting to ONNX?

Run a numerical comparison on at least 1000 samples from your validation set. Feed the same preprocessed input to the original framework and to ONNX Runtime (CPU provider, FP32) and compute cosine similarity or mean absolute error per output. I automate this in CI and fail the build if similarity is below 0.999 or if the top-1 prediction disagrees on more than 0.5% of samples. Then repeat the test with the quantized INT8 model and your target execution provider.

Related Articles

References & Standards: FreeRTOS Documentation · Zephyr Project Documentation · MQTT Specification