After deploying vibration-based monitoring systems across twelve different factory floors over the last five years, I've learned that predictive maintenance is rarely limited by the machine learning model itself. The real challenge lies in capturing clean, time-synchronized vibration data on resource-constrained hardware, processing it deterministically at the edge, and deciding what constitutes an anomaly when every motor has its own mechanical personality. A cloud-only approach simply cannot meet the latency or bandwidth requirements for a 25.6 kHz sampling pipeline running 24/7. In this article, I'll walk through the complete architecture I use for edge-native vibration analysis — from selecting the right MEMS accelerometer and designing the acquisition chain, to extracting robust spectral features and deploying a quantized autoencoder that runs in under 60ms on a Cortex-M4F. This is not theoretical; it's the stack that currently monitors over 400 rotating assets for my clients.
Decoding Vibration Signatures: From Time-Domain Waveforms to Bearing Fault Frequencies
Raw vibration data looks like noise until you understand the underlying physics. Every rotating asset — whether it's a 15kW induction motor, a gearbox, or a centrifugal pump — generates a deterministic set of frequencies related to its geometry and rotational speed. When a defect develops, it modulates these frequencies in predictable ways.
In my experience, the most reliable early indicator of failure is not overall RMS amplitude, but the emergence of specific fault frequencies. For rolling element bearings, you can calculate these directly:
Calculating Characteristic Bearing Frequencies
For a bearing with N rolling elements, pitch diameter D, ball diameter d, contact angle φ, and shaft speed fr (Hz), the four critical frequencies are:
BPFO (Outer Race): (N/2) * fr * (1 - (d/D) * cos φ)
BPFI (Inner Race): (N/2) * fr * (1 + (d/D) * cos φ)
BSF (Ball Spin): (D/(2*d)) * fr * (1 - ((d/D) * cos φ)^2)
FTF (Cage): (1/2) * fr * (1 - (d/D) * cos φ)
I've found that tracking the amplitude growth at BPFO and BPFI over weeks provides a far more stable health indicator than any broadband metric. A healthy bearing will show a flat spectrum with a clear 1x RPM peak and low harmonics. A bearing with an outer race spall will develop a sharp peak at BPFO with sidebands spaced at FTF. An imbalance, by contrast, shows a dominant 1x amplitude with little else. Misalignment typically drives up the 2x component.
Why Time-Domain Statistics Alone Will Fool You
Many early systems I audited relied solely on time-domain features like RMS, peak, crest factor, and kurtosis. Kurtosis is excellent for detecting impulsive impacts from early-stage spalling — a healthy signal has a kurtosis near 3.0, while a faulty one can spike above 6.0 — but it is extremely sensitive to load variations and transient shocks from normal operation. I once spent two weeks chasing false alarms on a stamping press line before realizing the piecewise impact loading was inflating kurtosis on every cycle. You need both domains. Time-domain features give you speed and low compute cost; frequency-domain features give you diagnostic specificity. A robust system fuses them.
Architecting the Edge Vibration Node: MEMS Accelerometers, Sampling Chains and Anti-Aliasing
The sensor selection defines the ceiling for your entire system. For predictive maintenance, you need bandwidth well beyond what consumer IMUs provide. Most bearing faults I monitor manifest between 500 Hz and 8 kHz, and for gear mesh frequencies that can extend to 10-12 kHz.
I now standardize on industrial MEMS accelerometers like the ADXL1002 or IIS3DWB, not general-purpose IMUs. The key specs I prioritize are: noise density below 25 µg/√Hz, bandwidth of at least 10 kHz, and a linear range of ±50g. The widely used MPU-6050 style parts top out at 1-2 kHz bandwidth and will alias or completely miss critical fault content. If you are combining vibration with orientation or speed data, refer to dedicated Sensor Fusion Algorithms: Combining IMU, GPS and Magnetometer Data for proper time alignment, but keep your vibration chain independent.
Sampling, Anti-Aliasing and Clock Discipline
To satisfy Nyquist for a 10 kHz signal of interest, you need at least 20 kSPS, but I sample at 25.6 kSPS. That number is not arbitrary — 25.6k gives exactly 25600 samples per second, which yields a clean 25 Hz bin resolution with a 1024-point FFT without spectral leakage windowing artifacts, and it aligns nicely with powers of two for DMA transfers.
You absolutely need an analog anti-aliasing filter before the ADC. I use a 2nd-order Sallen-Key low-pass at 11 kHz followed by oversampling and digital decimation. Without it, high-frequency energy from impacts folds back into your analysis band and creates phantom faults. I learned this the hard way on a high-speed spindle where a 14 kHz resonance aliased down to 3.2 kHz and triggered a misdiagnosis of a gear fault.
Clock discipline matters. Use a dedicated timer-triggered ADC with DMA, not a software-polling loop. Jitter in the sampling interval spreads spectral energy and raises your noise floor by 5-10 dB. On STM32 and nRF52 platforms, I drive the ADC from a hardware timer and double-buffer via DMA so the CPU only wakes when a full window is ready for processing.
Feature Engineering on Constrained Silicon: FFT, Spectral Kurtosis and Cepstral Pipelines
On a Cortex-M4F with 256KB RAM, you cannot run a full Python feature pipeline. You have to design for fixed-point or single-precision float, in-place computation, and minimal allocation. My standard edge pipeline takes a 2048-sample window (80ms at 25.6 kSPS), applies a Hanning window, computes a 2048-point real FFT via CMSIS-DSP, and extracts 32 features in under 35ms.
The Feature Set That Actually Survives Production
From the FFT magnitude spectrum, I extract: RMS in 8 logarithmically-spaced bands (10-100Hz, 100-300Hz, etc.), amplitudes at 1x, 2x, 3x RPM, BPFO/BPFI amplitudes, spectral centroid, spectral roll-off, and spectral kurtosis. From the time domain, I keep only RMS, kurtosis, and crest factor. That 32-dimensional vector is surprisingly robust across load conditions.
Spectral kurtosis deserves special attention. Unlike time-domain kurtosis, it identifies which frequency bands contain impulsive, non-stationary content — exactly what a localized bearing fault produces. The cepstrum (IFFT of log spectrum) is also invaluable for detecting families of harmonics from gearboxes without manually tracking each harmonic.
Here is the core of my CMSIS-DSP feature extraction loop that runs on-device:
#include "arm_math.h"
#include "arm_const_structs.h"
#define FFT_SIZE 2048
#define SAMPLE_RATE 25600.0f
float32_t window[FFT_SIZE];
float32_t fft_buf[FFT_SIZE * 2];
float32_t mag[FFT_SIZE/2];
void extract_vibration_features(float32_t *raw, float32_t *features) {
// Apply Hanning window in-place
arm_mult_f32(raw, window, fft_buf, FFT_SIZE);
// Real FFT using CMSIS-DSP
arm_rfft_fast_instance_f32 fft_inst;
arm_rfft_fast_init_f32(&fft_inst, FFT_SIZE);
arm_rfft_fast_f32(&fft_inst, fft_buf, fft_buf, 0);
// Magnitude: sqrt(re^2 + im^2)
arm_cmplx_mag_f32(fft_buf, mag, FFT_SIZE/2);
// Band energy features (example: 500-1500 Hz)
uint32_t band_start = (500 * FFT_SIZE) / (int)SAMPLE_RATE;
uint32_t band_end = (1500 * FFT_SIZE) / (int)SAMPLE_RATE;
float32_t band_energy;
arm_power_f32(&mag[band_start], band_end - band_start, &band_energy);
features[0] = sqrtf(band_energy);
// Spectral kurtosis simplified
float32_t mean, var, skew, kurt;
arm_mean_f32(mag, FFT_SIZE/2, &mean);
// ... variance/kurtosis calculation
// features[1..31] filled similarly
}
For preprocessing alternatives and how to handle quantization noise in this pipeline, the techniques in Edge AI Model Optimization: Pruning, Quantization and Knowledge Distillation are directly applicable, especially the section on per-channel quantization for spectral inputs.
Synchronization with RPM
If shaft speed varies, your frequency bins will smear. I always add a tachometer input — either a Hall sensor or an optical encoder — and use order tracking. In firmware, I resample the vibration data in the angular domain so that fault frequencies remain at fixed bins regardless of RPM. Without order tracking, variable-speed drives will destroy your model's accuracy during ramp-up and ramp-down phases.
Training and Compressing Anomaly Detection Models for Cortex-M and RISC-V Targets
Supervised fault classification is attractive in demos but impractical at scale. You will never have enough labeled examples of every failure mode for every asset class. I've shifted entirely to semi-supervised anomaly detection: train only on healthy data, detect deviations.
My workhorse is a deep autoencoder: input 32 features → 16 → 8 → 4 (bottleneck) → 8 → 16 → 32. Trained with MSE reconstruction loss on healthy data, it learns the manifold of normal operation. During inference, a high reconstruction error signals an anomaly. A threshold set at the 99th percentile of healthy reconstruction errors gives me a stable operating point. For rotating equipment, this outperforms isolation forests because it captures non-linear correlations between frequency bands.
Training happens in Python with TensorFlow/Keras, but deployment requires aggressive compression. A float32 autoencoder at 32→16→8 is already small, but after int8 quantization it fits in 18KB flash and runs in 12ms on an STM32F4 at 80MHz.
import tensorflow as tf
# Build autoencoder
inputs = tf.keras.Input(shape=(32,))
x = tf.keras.layers.Dense(16, activation='relu')(inputs)
x = tf.keras.layers.Dense(8, activation='relu')(x)
bottleneck = tf.keras.layers.Dense(4, activation='relu')(x)
x = tf.keras.layers.Dense(8, activation='relu')(bottleneck)
x = tf.keras.layers.Dense(16, activation='relu')(x)
outputs = tf.keras.layers.Dense(32)(x)
autoencoder = tf.keras.Model(inputs, outputs)
autoencoder.compile(optimizer='adam', loss='mse')
autoencoder.fit(healthy_features, healthy_features, epochs=80, batch_size=256, validation_split=0.1)
# Full integer quantization for TinyML
def representative_dataset():
for sample in healthy_features[:500]:
yield [sample.astype('float32').reshape(1, 32)]
converter = tf.lite.TFLiteConverter.from_keras_model(autoencoder)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_dataset
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_OPS_INT8]
converter.inference_input_type = tf.int8
converter.inference_output_type = tf.int8
tflite_model = converter.convert()
open("vibration_ae_int8.tflite", "wb").write(tflite_model)
The choice of toolchain depends heavily on your hardware target and framework preference. I evaluate them on every project:
| Deployment Stack | Model Format | Typical RAM / Flash | Strengths for Vibration Use Case |
|---|---|---|---|
| TensorFlow Lite for Microcontrollers | .tflite (int8) | 16-22 KB RAM / 28-45 KB Flash | Best tooling for autoencoders, excellent CMSIS-NN kernels, easy quantization |
| ONNX Runtime (ARM Build) | .onnx (int8/fp16) | 28-40 KB RAM / 60-90 KB Flash | Framework-agnostic, ideal if training in PyTorch, good for heterogeneous fleets |
| Custom CMSIS-DSP / C Array | C header (int8/int16) | 8-12 KB RAM / 15-25 KB Flash | Minimal overhead, deterministic, but manual quantization and no graph optimizations |
In practice, I deploy TFLite Micro for Cortex-M projects and keep ONNX as my fallback for mixed architectures. If you are standardizing on ONNX for cross-platform reuse, see ONNX Runtime for Edge: Cross-Framework Model Deployment on ARM for ARM-specific build flags and operator coverage. For a deeper walkthrough of running these quantized models on microcontrollers, TinyML: Running Neural Networks on Microcontrollers with TensorFlow Lite provides the memory planning workflow I still use.
Streaming vs. Store-and-Forward: MQTT, Ring Buffers and Deterministic Inference Scheduling on FreeRTOS
You cannot stream 25.6 kSPS raw data over MQTT or BLE continuously — a single node would saturate the network. The edge must decide what to send. My nodes operate on a hierarchical reporting strategy: inference runs every 2 seconds locally, and only feature vectors, anomaly scores, and occasional raw windows are published.
FreeRTOS gives me the determinism I need. I structure the firmware as three tasks with strict priorities, as recommended in the FreeRTOS Documentation:
1. Acquisition Task (Highest Priority): Handles DMA completion interrupts, copies the double buffer, and pushes to a lock-free ring buffer. Never blocks.
2. Inference Task (Medium Priority): Pops a window from the ring buffer, runs windowing, FFT, and autoencoder inference. Runs every 2048 samples.
3. Comms Task (Lowest Priority): Publishes results via MQTT. Packets are queued; if connectivity drops, data is spooled to external flash.
This separation prevents network jitter from causing sample drops. I've seen designs where MQTT publishing blocked the sampling loop — the resulting clock drift alone increased false positives by 18%.
MQTT Topic and Payload Design
I use a topic structure like factory/line3/motor-07/vibration/anomaly with QoS 1 for anomaly scores and QoS 0 for periodic health heartbeats. Payloads are CBOR-encoded, not JSON, to save 40% bandwidth. Raw waveforms are only sent on-demand via an MQTT request-response pattern, following the MQTT Specification for retained messages and clean session handling. A retained message on .../status ensures a new dashboard client immediately knows if a node is offline.
// FreeRTOS inference task pseudocode
void vInferenceTask(void *pvParameters) {
float raw_window[2048];
float features[32];
for(;;) {
// Wait for buffer ready notification from ISR
xTaskNotifyWait(0, 0, NULL, portMAX_DELAY);
if(xRingBufferReceive(vibRingBuf, raw_window, sizeof(raw_window), 0) == pdTRUE) {
extract_vibration_features(raw_window, features);
// TFLite Micro inference
TfLiteTensor* input = interpreter->input(0);
quantize_input(features, input->data.int8, 32);
interpreter->Invoke();
float anomaly_score = dequantize_and_compute_mse(interpreter->output(0));
if(anomaly_score > ANOMALY_THRESHOLD) {
// Publish score + feature vector, optionally trigger raw capture
xQueueSend(mqttQueue, &anomaly_msg, 0);
// Also set retained status
}
}
}
}
If you are managing a fleet of these nodes with cloud synchronization and OTA model updates, an orchestrator like AWS IoT Greengrass: Edge Computing, Local Processing and Cloud Sync handles the shadow sync and deployment mechanics so you don't have to build it from scratch. I've used Greengrass to roll out a new quantized model to 80 nodes in under an hour with automatic rollback on inference errors.
Field Deployment Realities: Calibration Drift, Temperature Compensation and False Positive Suppression
Lab accuracy rarely survives the factory floor. The two issues that consistently degrade performance after deployment are sensor drift and operating context changes.
MEMS accelerometers drift with temperature — sensitivity can shift 0.02% per degree Celsius. On a motor housing that cycles from 25°C to 75°C, that is a 1% gain error, enough to push a borderline anomaly score over threshold. I always mount a temperature sensor (TMP117 or the accelerometer's internal temp sensor) and apply a linear compensation to the raw ADC values before feature extraction. Calibration should be performed after thermal soak, not on a cold bench.
Mounting is even more critical. A loosely torqued stud or a magnet mount with a thin oil film introduces a resonance around 2-4 kHz that looks identical to a gear fault. I specify stud mounting with 2.5 N·m torque and a thin layer of silicone grease, and I verify mounting resonance with a bump test during commissioning. Adhesive mounts are acceptable only below 5 kHz analysis bands.
Suppressing False Positives Without Missing Failures
An anomaly score that spikes once is not actionable — every load change or tool engagement will cause a transient. I apply three layers of filtering:
1. Median filtering over a 30-second window: This removes single-window spikes while preserving sustained fault signatures.
2. Hysteresis thresholds: Alert only when the median exceeds 1.4x the healthy baseline for 5 consecutive inference cycles, clear only when it drops below 1.2x for 10 cycles. This prevents flapping.
3. Context gating: I gate inference on steady-state RPM. If tachometer variance exceeds 5% within the window, I discard the window. This alone cut false positives by 60% on variable-frequency drive systems in my last deployment.
Finally, log everything locally. When a customer reports a missed failure, the first thing I pull is the last 100 feature vectors from flash. In one case, that log revealed that a new production recipe had shifted the baseline load profile — requiring a one-shot re-normalization of the model inputs. Predictive maintenance at the edge is not a deploy-and-forget system; it requires continuous baseline management, but when engineered correctly, it catches bearing failures 3-6 weeks before audible symptoms with less than 2% false alarm rate.
Frequently Asked Questions
What sampling rate is actually required for bearing fault detection?
For most industrial bearings with shaft speeds between 1000-3600 RPM, fault frequencies fall between 30 Hz and 6 kHz. To capture harmonics and early-stage impacts, I recommend a usable bandwidth of at least 10 kHz, which requires 25.6 kSPS sampling after proper analog anti-aliasing. Sampling at 1-5 kSPS, as many low-cost IoT vibration sensors do, will miss outer race and ball spin faults entirely.
Can I use a generic IMU like the MPU-6050 for predictive maintenance?
Not for reliable bearing diagnostics. Low-cost IMUs have bandwidth limits around 1 kHz, higher noise density (300-400 µg/√Hz), and internal digital low-pass filters that cannot be bypassed. They are suitable for imbalance or misalignment detection on low-speed machines but lack the sensitivity for incipient bearing faults. For production use, choose an industrial vibration MEMS with <25 µg/√Hz noise density and >5 kHz bandwidth.
How much labeled failure data do I need to train a TinyML model?
If you use a semi-supervised autoencoder approach, you need zero failure examples — only 2-4 hours of healthy data per asset class covering different loads and speeds to learn the normal manifold. I typically collect 20,000-30,000 windows of healthy data per machine type. Supervised classifiers require dozens of labeled failures per fault type, which is rarely feasible without accelerated life testing.
How do I handle model updates for nodes already deployed in the field?
Use delta OTA updates via your edge orchestrator and always maintain a fallback model in flash. I version both the model and the feature extraction code together, since a quantization parameter change affects thresholds. Before activating a new model, the node runs it in shadow mode for 24 hours, comparing its scores to the active model without triggering alerts, then switches over only if drift is within tolerance.