Raspberry Pi was never designed for the factory floor, yet I keep deploying it there. Over the last five years I have installed more than forty Pi-based gateways in small manufacturing cells, water pumping stations, and building management retrofits where a full PLC or industrial PC would have blown the budget. The appeal is obvious: for $45-$75 you get a quad-core ARM board with Ethernet, serial, and enough headroom to translate legacy fieldbus to modern cloud protocols. The challenge is making that consumer board behave like reliable infrastructure. In this article I will walk through how I architect a raspberry pi as a robust iot gateway that polls modbus devices, normalizes data at the edge, and publishes to MQTT without losing messages when the network or power inevitably falters. This is not a lab demo; it is the pattern I use in production.
Why the Raspberry Pi Survives on the Factory Floor Despite Not Being Industrial-Grade
The first question every controls engineer asks me is why not use an ESP32 or a ruggedized gateway. The answer comes down to protocol translation overhead. A typical cell I worked with recently had eight Modbus RTU energy meters daisy-chained on RS-485, two Modbus TCP PLCs, and a requirement to push JSON to AWS IoT Core over TLS. An ESP32 for IoT: WiFi, Bluetooth and Deep Sleep for Battery-Powered Projects is excellent for battery-powered sensor nodes, but it struggles when you need concurrent TCP sockets, TLS, and a Python or Node-RED runtime for data mapping. The Pi gives you a full Debian userland, which matters when you need to run OpenSSL, buffering databases, and remote management agents.
In my experience, the Pi 4 Model B and the Compute Module 4 are the only variants worth considering for industrial use today. The Pi 3 is underpowered for TLS at scale, and the Zero lacks Ethernet. The Compute Module 4 is my preference for anything permanent. It exposes native PCIe, allows you to design a carrier board with proper RS-485 transceivers, TVS diodes, and a wide-input 9-36V buck converter. You also get eMMC, which is far more reliable than an SD card. If you must use a standard Pi 4, budget for an industrial SD card rated for high endurance and power-loss protection, or better, boot from an external SSD.
Choosing Between Pi 4, Compute Module 4, and True Industrial Hardware
Thermal and power design will determine your uptime more than software. I've found that a Pi 4 in a sealed ABS enclosure without a heatsink will throttle at 80°C within 30 minutes in a 35°C electrical cabinet. Use a DIN-rail enclosure with an aluminum heatsink case acting as thermal mass, or add a 40mm fan with a simple hysteresis controller. For power, never rely on a consumer USB-C wall adapter. A Mean Well HDR-30-24 plus a 24V to 5V 3A buck with reverse polarity and surge protection will keep the Pi alive through brownouts that reset cheaper supplies. The extra $40 saves a site visit.
The SD Card Problem and How I Solve It
SD card corruption after an unplanned power cycle is the classic Pi failure. I run Raspberry Pi OS Lite (64-bit) with the root filesystem mounted read-only and overlayfs enabled via raspi-config. All writable data goes to a separate partition or an external USB SSD, and logs are shipped via journald to a remote syslog instead of written continuously. This single change has taken my field failure rate from roughly one corruption per six months to zero in two years across 22 nodes.
Wiring RS-485 and Polling Modbus RTU/TCP Without Blocking Your Gateway
Most legacy equipment still speaks Modbus RTU over RS-485. The Pi has no native RS-485 interface, so the transceiver choice matters. I avoid cheap USB-to-RS-485 dongles based on CH340 for production; their driver latency and lack of automatic direction control cause framing errors at 38400 baud and above. For a Pi 4, I use a Waveshare RS485 CAN HAT that uses the SP3485 with hardware DE/RE control tied to GPIO, or on a Compute Module carrier, a proper isolated ADM2483 transceiver with 2.5kV isolation and 120-ohm termination selectable via jumper.
Wiring discipline is critical. RS-485 is a differential bus that requires daisy-chaining, not a star topology. Keep A/B twisted, add 120-ohm termination only at the two physical ends of the bus, and add 680-ohm bias resistors if your transceiver does not provide them. I have debugged more than one unstable bus that turned out to be a floating ground; tie the Pi's RS-485 GND to the Modbus device common through a 100-ohm resistor to avoid ground loops while providing a reference.
Non-Blocking Modbus Polling in Python
For polling I use pymodbus 3.x, which finally cleaned up its async support. Do not poll in a tight loop on the main thread. I run a dedicated Modbus thread that polls on a fixed schedule and pushes decoded register values into a thread-safe queue. This decouples bus timing from MQTT publishing. Below is the core pattern I use for Modbus RTU polling with automatic reconnection:
from pymodbus.client import ModbusSerialClient
from pymodbus.exceptions import ModbusIOException
import time
from queue import Queue
queue = Queue(maxsize=1000)
client = ModbusSerialClient(
port='/dev/ttySC0',
baudrate=19200,
bytesize=8,
parity='E',
stopbits=1,
timeout=1.0,
retries=1
)
REGISTER_MAP = {
"energy_meter_1": {"slave": 1, "address": 0x0000, "count": 10},
"energy_meter_2": {"slave": 2, "address": 0x0000, "count": 10},
}
def poll_loop():
if not client.connect():
print("Modbus serial connect failed")
time.sleep(5)
return
while True:
cycle_start = time.monotonic()
for name, cfg in REGISTER_MAP.items():
try:
rr = client.read_holding_registers(
address=cfg["address"],
count=cfg["count"],
slave=cfg["slave"]
)
if not rr.isError():
# decode registers to engineering units
payload = decode_registers(name, rr.registers)
if not queue.full():
queue.put((name, payload))
else:
print(f"Modbus error {name}: {rr}")
except ModbusIOException as e:
print(f"Bus exception {name}: {e}")
time.sleep(0.5)
# maintain fixed 1s cycle time regardless of poll duration
elapsed = time.monotonic() - cycle_start
time.sleep(max(0, 1.0 - elapsed))
For Modbus TCP devices on the same gateway, I instantiate a second client using ModbusTcpClient with a 2-second timeout. I never mix RTU and TCP on the same client instance. Keep TCP polling in a separate thread as well; a stalled TCP socket should not block your RTU bus.
Register Mapping and Scaling Lessons
Document your register map in a single YAML file versioned in Git. I include slave ID, function code, data type (uint16, int32, float32, with ABCD/CDAB endian), scale factor, and units. I've wasted hours because a vendor documented 32-bit energy values as big-endian when the device actually sent little-endian. Always validate with a known load. Tools like modpoll or CAS Modbus Scanner are useful for initial discovery before you write any gateway code.
Bridging Modbus Registers to MQTT Topics: Payload Design That Scales
The gateway's main job is protocol translation: Modbus is poll-based and register-oriented, MQTT is publish-subscribe and message-oriented. How you bridge those models determines whether your cloud application can scale. I follow the Sparkplug B-influenced topic structure even when not strictly using Sparkplug: site/area/device/metric. For example, plant1/line2/meter01/active_power is far more useful than dumping all registers into one JSON blob per device. It allows fine-grained subscriptions and avoids parsing overhead downstream.
I use Eclipse Paho MQTT Python client with TLS mutual authentication and offline buffering. The MQTT Specification (https://mqtt.org/mqtt-specification/) defines QoS 0, 1, and 2, but in practice for sensor telemetry I use QoS 1 with persistent sessions and clean_session=False. QoS 2 adds too much overhead for little benefit on stable TCP, and QoS 0 will silently drop data during reconnections.
import paho.mqtt.client as mqtt
import json, time, ssl
mqtt_client = mqtt.Client(client_id="pi-gateway-line2-01", clean_session=False)
mqtt_client.tls_set(
ca_certs="/etc/gateway/certs/ca.crt",
certfile="/etc/gateway/certs/gateway.crt",
keyfile="/etc/gateway/certs/gateway.key",
tls_version=ssl.PROTOCOL_TLSv1_2
)
mqtt_client.username_pw_set("gateway", "strong-password-here")
mqtt_client.will_set("plant1/line2/gateway01/status", "offline", qos=1, retain=True)
mqtt_client.reconnect_delay_set(min_delay=1, max_delay=60)
def on_connect(client, userdata, flags, rc):
if rc == 0:
client.publish("plant1/line2/gateway01/status", "online", qos=1, retain=True)
mqtt_client.on_connect = on_connect
mqtt_client.connect("mqtt.example.com", 8883, 60)
mqtt_client.loop_start()
# Bridge from Modbus queue to MQTT
while True:
name, data = queue.get()
topic = f"plant1/line2/{name}/telemetry"
payload = json.dumps({"ts": int(time.time()), "values": data})
# QoS 1, retain False for telemetry
mqtt_client.publish(topic, payload, qos=1, retain=False)
Notice the Last Will and Testament on the status topic. This gives any subscriber immediate visibility into gateway offline events without a separate heartbeat. I also publish a retained birth certificate with firmware version and register map hash so the backend can detect configuration drift.
Payload Normalization at the Edge
Never publish raw register values. Convert to engineering units on the gateway: scale a 0-10000 raw value to 0-100.0 kW, apply signed conversion, and timestamp at sampling time, not publish time. This offloads conversion from the cloud and makes the data self-describing. I keep a small Python module that takes the YAML register map and generates decode functions automatically, which eliminates manual scaling errors.
Edge Processing on ARM: Filtering, Buffering and Local Decision Logic
Publishing every Modbus sample to the cloud is wasteful. A Pi 4 can easily handle edge processing that reduces bandwidth by 80% and adds resilience when the WAN link drops. In my deployment for a cold storage facility, the gateway samples temperature every second but only publishes on a 5-second interval or immediately if the value deviates by more than 0.5°C. This deadband filtering cut MQTT traffic from 86,400 messages per day per sensor to under 15,000 with no loss of critical information.
Buffering is the other critical edge function. I use a local SQLite database or a persistent Paho queue to store messages when the broker is unreachable. The pattern is simple: if publish() returns MQTT_ERR_NO_CONN or the client is not connected, write to SQLite. A separate flusher thread retries with exponential backoff when connectivity returns, preserving order via timestamps. For time-series data, SQLite with WAL mode handles thousands of inserts per second on a Pi 4 without wearing the storage excessively.
import sqlite3
import time
db = sqlite3.connect("/data/buffer.db", check_same_thread=False)
db.execute("CREATE TABLE IF NOT EXISTS outbox (ts INTEGER, topic TEXT, payload TEXT)")
def buffered_publish(topic, payload):
if mqtt_client.is_connected():
info = mqtt_client.publish(topic, payload, qos=1)
if info.rc != mqtt.MQTT_ERR_SUCCESS:
db.execute("INSERT INTO outbox VALUES (?,?,?)", (int(time.time()), topic, payload))
db.commit()
else:
db.execute("INSERT INTO outbox VALUES (?,?,?)", (int(time.time()), topic, payload))
db.commit()
def flush_loop():
while True:
if mqtt_client.is_connected():
cur = db.execute("SELECT rowid, topic, payload FROM outbox ORDER BY ts LIMIT 50")
rows = cur.fetchall()
for rowid, topic, payload in rows:
mqtt_client.publish(topic, payload, qos=1).wait_for_publish(timeout=2)
db.execute("DELETE FROM outbox WHERE rowid=?", (rowid,))
db.commit()
time.sleep(5)
Local decision logic is where the gateway justifies its existence. I implement simple threshold alarms directly in Python: if a Modbus pressure reading exceeds 7 bar for 10 seconds, the gateway can drive a GPIO relay or write back to a Modbus coil to shut a valve, without waiting for a cloud round-trip. This requires careful state management. I use a finite state machine with debounce timers to avoid chattering, and I log every automated action to a local audit table for later review.
When to Use Node-RED vs. Custom Code
Node-RED is attractive for rapid prototyping and I still use it for internal tools, but I have moved production gateways to plain Python services. Node-RED flows become difficult to version, test, and monitor once they exceed 20 nodes. A systemd-managed Python package with unit tests and structured logging is more maintainable for a gateway that must run for years. If your team is not comfortable with Python, consider a compiled Go service for even lower memory footprint.
Hardening the Pi Gateway: Watchdogs, Read-Only Filesystems and Power Resilience
Reliability engineering separates a hobby project from an industrial gateway. I harden every Pi with four layers. First, the hardware watchdog. The BCM2711 has a built-in watchdog that I enable via dtparam=watchdog=on in config.txt and configure with watchdog or systemd. If the Python process deadlocks or the kernel panics, the board resets within 15 seconds. Second, systemd service hardening: set Restart=always, RestartSec=5s, and StartLimitBurst=5 for the gateway service, and use WatchdogSec=30s with sd_notify from Python so a hung poll loop triggers a restart.
Third, power resilience. Even with a good supply, I add a small UPS. For DIN-rail installations I use a Phoenix Contact QUINT-UPS or a simple supercapacitor UPS HAT that gives 20-30 seconds to flush buffers and shut down cleanly. The gateway monitors the UPS status via GPIO and initiates sync and service stop when mains fail. Without this, even a read-only filesystem can lose buffered data in RAM.
Fourth, remote observability. Each gateway runs node_exporter for Prometheus and a lightweight WireGuard VPN for SSH access. I expose metrics like Modbus poll latency, MQTT queue depth, CPU temperature, and buffer row count. Alerting on buffer growth has caught upstream broker issues before the customer noticed. In my experience, investing a day in observability saves a week of on-site debugging.
| Gateway Option | Typical Cost | Modbus Handling | Edge Compute Headroom | Industrial Hardening Effort |
|---|---|---|---|---|
| Raspberry Pi 4 (with HAT) | $85 - $140 | RS-485 via HAT + Modbus TCP via Ethernet | High: Python, Docker, SQLite, TLS | Medium: requires UPS, cooling, read-only FS |
| Raspberry Pi Compute Module 4 | $120 - $180 | Isolated RS-485 on carrier, dual Ethernet | High: same as Pi 4, eMMC more reliable | Low-Medium: carrier can include isolation/UPS |
| ESP32 Gateway | $15 - $35 | Modbus RTU via UART, limited TCP | Low: C/C++, limited TLS concurrency | Low: inherently more robust, less to corrupt |
| Industrial PC / PLC Gateway | $400 - $1200 | Native isolated RS-485 / PROFINET | High: x86, often Windows/Linux | Minimal: DIN-rail, 24V, certified |
This comparison guides my Microcontroller Selection: ARM Cortex-M, RISC-V and AVR Decision Criteria discussions with clients. If the site has fewer than 15 Modbus devices and no complex edge logic, an ESP32 can be sufficient and more robust. Once you need TLS client certificates, local buffering, or multi-protocol translation, the Pi's Linux environment pays for itself. For harsh environments above 60°C or with strict certification requirements, I still recommend a certified industrial gateway and reserve the Pi for non-critical monitoring.
Field-Tested Software Stack: From systemd Services to Docker on Raspberry Pi OS
My current production image is Raspberry Pi OS Lite 64-bit, provisioned via cloud-init or Ansible. The gateway application is packaged as a Python wheel with dependencies pinned in requirements.txt. I avoid Docker on most gateways to keep overhead low, but for gateways that run multiple services like Node-RED, Mosquitto bridge, and the Modbus poller, Docker Compose simplifies updates. If you use Docker, pin the base image to python:3.11-slim-bullseye and set memory limits to prevent one container from starving the system.
Each gateway service is a systemd unit with hardened directories. Here is a trimmed unit I use:
[Unit]
Description=Modbus to MQTT Gateway
After=network-online.target
Wants=network-online.target
[Service]
Type=notify
ExecStart=/opt/gateway/venv/bin/python -m gateway.main
Restart=always
RestartSec=5
WatchdogSec=30
User=gateway
WorkingDirectory=/opt/gateway
Environment=PYTHONUNBUFFERED=1
# Hardening
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ReadWritePaths=/data /var/log/gateway
[Install]
WantedBy=multi-user.target
OTA updates are handled via Mender or a simple A/B root partition scheme with rauc. I never allow apt upgrade unattended; a kernel update that changes the serial driver can break RS-485 timing. Instead, I build a golden image, test it on a bench Pi connected to a Modbus simulator like diagslave, and push it as a full image update.
Security Baseline You Should Not Skip
Change the default pi password, disable password SSH and use keys only, enable UFW with only 22 and 8883 outbound, and store private keys with 0600 permissions. For MQTT, use per-gateway certificates issued by your private CA rather than shared credentials. I also isolate the Modbus network physically; the RS-485 bus should never be bridged to the corporate LAN except through the gateway's application logic. Developers coming from STM32 HAL Drivers: GPIO, UART, SPI and DMA Configuration backgrounds often expect to bit-bang everything, but on the Pi you get to rely on the kernel's serial driver and focus your effort on application reliability and security.
Logging is structured JSON to journald with a forwarder to Loki or CloudWatch. Include slave ID, function code, latency, and error codes in every Modbus log line. When a site reports gaps in data, those logs let me distinguish between a bus wiring fault and a broker outage within minutes. The combination of structured code, robust OS configuration, and edge buffering is what lets a $50 board meet the reliability expectations of industrial monitoring.
Frequently Asked Questions
Can a Raspberry Pi reliably handle both Modbus RTU and Modbus TCP concurrently?
Yes, if you isolate them. Run separate client instances and threads for RTU and TCP. The RTU bus is sensitive to timing jitter, so keep its poll loop dedicated and avoid blocking calls inside it. TCP clients can tolerate more latency but need independent timeouts. I routinely poll one RTU bus at 1 second and two TCP PLCs at 500 ms on a Pi 4 without issues, but I monitor poll loop duration to ensure it stays well below the cycle time.
How do you prevent SD card failure in continuous operation?
Use a high-endurance card, enable read-only root with overlayfs, and redirect all writes to a separate data partition on eMMC or SSD. Disable swap, limit log writes, and add a small UPS to allow graceful flush on power loss. For permanent installations, I strongly recommend the Compute Module 4 with eMMC, which eliminates the SD card entirely.
Should the gateway translate Modbus registers to JSON or use a binary format like Sparkplug B?
For fewer than 200 data points, plain JSON over MQTT with a well-structured topic hierarchy is simpler to debug and integrates with most cloud platforms. Sparkplug B adds efficient binary encoding and state management but requires a compatible broker and client libraries. I start with JSON and migrate to Sparkplug or Protobuf only when bandwidth or payload size becomes a measured bottleneck.
Is Docker recommended for a Pi-based IoT gateway?
Docker helps when you need multiple isolated services, but it adds CPU and memory overhead and complicates access to serial hardware. For a single-purpose Modbus-to-MQTT gateway I deploy a native systemd service. If you need Mosquitto, Telegraf, and your gateway app on the same Pi, Docker Compose is justified. In either case, set resource limits and test serial device passthrough thoroughly.