MQTT has become the default messaging fabric for most of the IoT deployments I've built over the last eight years, from battery-powered sensor nodes reporting once an hour to industrial gateways pushing thousands of messages per second. Its strength is not just that it is lightweight, but that it gives you explicit control over delivery guarantees, state retention, and session continuity — three areas where embedded systems either succeed or fail silently in the field. When you move beyond a simple publish and subscribe demo, you quickly find that understanding QoS, retained messages, and session management is essential to building a system that behaves predictably after a power loss, a cellular dropout, or a firmware update. In this article I will break down how each of these mechanisms works at the protocol level, how they interact, and how I have learned to apply them on constrained hardware.
QoS 0, 1, and 2 Under the Hood: Packet Flows and Delivery Guarantees
Quality of Service in MQTT is often taught as three simple levels, but the implications for bandwidth, memory, and reliability are significant. As an iot protocol built on a pub sub model, MQTT decouples publishers from subscribers and lets the broker mediate delivery. QoS defines the contract between each client and the broker, not end-to-end between publisher and subscriber. A publisher can send at QoS 1 while a subscriber receives at QoS 0, and the broker handles the translation.
QoS 0: At Most Once and Why It Still Matters
QoS 0 is fire-and-forget. The PUBLISH packet is sent and no acknowledgment is expected. There is no retransmission, no storage on the broker for offline clients, and no duplicate handling. I use QoS 0 for high-frequency telemetry where losing a single sample does not affect the system — for example, temperature readings every 30 seconds where the next reading makes the last one obsolete. It minimizes packet overhead to just two bytes for the fixed header plus topic and payload, which matters on a narrowband link.
QoS 1: At Least Once With Duplicate Handling
With QoS 1, the sender stores the message until it receives a PUBACK from the receiver. If the PUBACK does not arrive within a timeout, the sender retransmits with the DUP flag set. This guarantees delivery but can produce duplicates if the PUBACK was lost. In my experience, most application bugs with mqtt come from assuming QoS 1 is exactly-once. Your subscriber code must be idempotent. I've found that adding a simple message ID or timestamp in the payload and de-duplicating in the consumer avoids double-processing actuator commands.
// Publishing with QoS 1 using Eclipse Paho Python
// Message will be retried until PUBACK is received
client.publish(
topic="factory/line1/vibration",
payload=json.dumps({"sensor_id": "vib-04", "rms": 3.42}),
qos=1,
retain=False
)
QoS 2: Exactly Once Through a Four-Way Handshake
QoS 2 is the most expensive and least used level, but it is critical when duplicate delivery would cause real harm, such as a payment or a one-time door unlock. It uses a four-step exchange: PUBLISH -> PUBREC -> PUBREL -> PUBCOMP. Both sides store the packet identifier until the handshake completes. The MQTT Specification requires that the receiver stores the identifier to discard any retransmitted PUBLISH with the same ID. On constrained devices, this means RAM for in-flight messages. I have seen ESP32-class devices run out of heap when 20 QoS 2 messages are queued during a cellular outage. Use it sparingly and always bound your in-flight window.
/* ESP-IDF / FreeRTOS MQTT publish with QoS 2 */
esp_mqtt_client_publish(client,
"access/door/main",
"unlock:nonce-9f3a",
0, // len (0 = auto)
2, // qos
0); // retain
// Ensure mqtt_cfg.buffer_size and task stack account for PUBREC/PUBREL handling
// See FreeRTOS Documentation for memory sizing guidance
| Feature | QoS 0 (At Most Once) | QoS 1 (At Least Once) | QoS 2 (Exactly Once) |
|---|---|---|---|
| Delivery Guarantee | Best effort, no Ack | Guaranteed, duplicates possible | Guaranteed, no duplicates |
| Packet Flow | PUBLISH | PUBLISH → PUBACK | PUBLISH → PUBREC → PUBREL → PUBCOMP |
| Sender Storage Required | None | Until PUBACK received | Until PUBCOMP received |
| Overhead & Latency | Minimal | Moderate, one extra packet | Highest, two extra round trips |
| Typical Use Case | Periodic telemetry, logs | Alerts, state updates requiring Ack | Critical commands, financial transactions |
One detail often missed: QoS is negotiated hop-by-hop. If you publish at QoS 2 but the subscriber subscribed at QoS 0, the subscriber will receive at most QoS 0. The broker downgrades. If you need end-to-end exactly-once, both sides must request QoS 2. I always document this explicitly in our topic architecture to avoid false assumptions.
Retained Messages as Distributed State: When and How to Use Them Safely
By default, MQTT is event-driven — if you subscribe after a message was published, you miss it. Retained messages solve this by turning the broker into a last-value cache. When a publisher sets the retain flag to 1, the broker stores the last retained message for that topic and immediately delivers it to any new subscriber. This is not a queue; only one message per topic is retained, and a new retained publish overwrites the old one.
The Retained Message Pattern for Configuration and Status
In my experience, retained messages are ideal for slow-moving state that new clients need instantly: device configuration, last reported status, or firmware version. For example, publishing device/thermostat-12/config as retained lets a mobile app get the current setpoint immediately on subscribe without polling or waiting for the device to republish. I've found that using retained messages for command topics is dangerous, because a device that reboots and resubscribes will immediately re-execute the last command.
# Python: publishing a retained configuration and clearing it
# Publish retained state
client.publish("building/floor2/thermostat-12/config",
payload='{"setpoint": 21.5, "mode": "auto"}',
qos=1, retain=True)
# To clear the retained message, publish a zero-byte payload with retain=True
client.publish("building/floor2/thermostat-12/config",
payload=None, qos=1, retain=True)
Pitfalls I Have Seen in Production
First, retained messages persist until overwritten or explicitly cleared, even across broker restarts if persistence is enabled. During development, stale retained messages cause confusing behavior — a node subscribes and gets a months-old value. I now include a version or timestamp in every retained payload and have a startup script that clears deprecated topics. Second, retained messages are per-topic, not per-client. If you use wildcards, you will receive one retained message per matching topic, which can create a burst of messages on subscribe. On a constrained gateway subscribing to sensors/# with 500 retained topics, that burst can overwhelm the receive buffer. Plan your topic tree so wildcard subscriptions do not unintentionally pull hundreds of retained payloads at once.
For contrast with protocols that handle state differently, it helps to compare with how LoRaWAN Network Architecture: Gateways, Network Server and Join Procedures manages device context at the network server, or how BLE exposes state through GATT characteristics.
Clean Start, Persistent Sessions and MQTT 5 Session Expiry
Session management determines what the broker remembers about a client after it disconnects. This is where MQTT 3.1.1 and MQTT 5 differ substantially, and where many field failures originate.
MQTT 3.1.1: Clean Session Flag
In MQTT 3.1.1, the CONNECT packet contains a Clean Session flag. If set to 1, the broker discards any previous session for that Client ID — subscriptions, queued QoS 1/2 messages, and pending acknowledgments. If set to 0, the broker creates a persistent session. Queued messages for that client are stored while it is offline and delivered on reconnection, respecting the subscription QoS.
The catch is that persistent sessions require a stable Client ID. I've seen fleets where every device used client_id = "sensor" and kept stealing each other's sessions. Use a unique, deterministic ID like esp32-24:6f:28:ab:12:cd or a provisioned serial number stored in NVS. Also, persistent sessions consume broker resources. A broker like Mosquitto or EMQX will hold queued messages until the client returns or until limits are hit. Without limits, an offline client subscribed to a high-throughput topic can cause memory exhaustion on the broker.
MQTT 5: Clean Start and Session Expiry Interval
MQTT 5 replaces Clean Session with two fields: Clean Start and Session Expiry Interval. Clean Start controls whether to start with a fresh session at connection time, while Session Expiry Interval in seconds tells the broker how long to keep the session after disconnection. A value of 0 means no persistence, 0xFFFFFFFF means infinite, and any other value is a timed persistence.
This is a major improvement for embedded systems. In my experience, setting Session Expiry to 3600 seconds (one hour) for battery devices that sleep is far safer than infinite persistence. If a device is decommissioned or reprovisioned, its session does not linger forever. I typically set Clean Start = 0 and Session Expiry = 86400 for mains-powered gateways, and Clean Start = 1 for ephemeral diagnostic tools that should never queue messages.
When implementing reconnection logic on Zephyr or FreeRTOS, I check the CONNACK Session Present flag. If Session Present is 0 when I expected a resumed session, I know to resubscribe. According to the Zephyr Project Documentation, the Zephyr MQTT stack exposes this flag directly in the connect ACK event, which makes it straightforward to implement conditional resubscribe logic without redundant SUBSCRIBE packets.
Last Will, Keep-Alives and Detecting Ungraceful Client Failures
Because MQTT runs over TCP, a client that loses power or cellular coverage does not send a DISCONNECT. The broker needs other mechanisms to infer liveness. Two are essential: Keep Alive and Last Will and Testament (LWT).
Keep Alive and Its Interaction With Sleep
Keep Alive is the interval in seconds declared in CONNECT. The client must send a PINGREQ if no other packet was sent within 1.5x that interval; otherwise the broker closes the connection. For always-on devices, I use 60 seconds. For deep-sleep nodes that wake every 10 minutes, I set Keep Alive to 300 seconds and disable PINGREQ during sleep — the session expiry and LWT handle the offline detection instead. Setting Keep Alive too low on a congested NB-IoT link causes false disconnects; I've found that 120 seconds is a safer default on cellular than the 15 seconds often used on Wi-Fi.
Designing a Useful Last Will and Testament
LWT is a message the client registers at connect time. If the broker detects an ungraceful disconnect (Keep Alive timeout or TCP close without DISCONNECT), it publishes the will message on the client's behalf. In MQTT 5, you can also set Will Delay Interval so the will is not published immediately — useful to tolerate brief glitches.
I always set a will for status topics, not for command topics. A good pattern is to publish device/ as retained with payload {"state":"online"} on connect, and register a will with the same topic, retained, payload {"state":"offline"} and QoS 1. That way, any subscriber instantly knows the device's liveness without polling. In one factory deployment, we added reason and timestamp fields to the will payload after we realized we could not distinguish a power cut from a firmware crash. That small addition cut our debug time in half.
If your system uses BLE for provisioning and MQTT for backhaul, understanding both session models helps. The connection supervision timeout in BLE Development: GATT Services, Advertising and Connection Management is conceptually similar to Keep Alive, though BLE manages it at the link layer.
Selecting QoS Levels and Session Strategy for Constrained Deployments
Choosing QoS and session settings is not about picking the highest guarantee everywhere. It is about matching the guarantee to the consequence of loss, duplication, and broker storage. On a device with 80 KB RAM, every queued QoS 1 message can be a liability during an outage.
My Decision Framework After Several Field Rollouts
I categorize messages into three classes. Telemetry class: sensor readings that are frequent and idempotent. Use QoS 0, retain false, clean session is irrelevant because delivery is best effort. State class: door is open/closed, thermostat setpoint, occupancy. Use QoS 1, retain true, persistent session with a bounded expiry so late joiners get the correct state. Command class: unlock, reboot, transfer funds. Use QoS 2, retain false, persistent session only if the command must survive offline periods, otherwise ephemeral.
For subscriptions, I avoid QoS 2 on the gateway unless the command is critical. Subscribing at QoS 1 to a state topic with retained messages and a persistent session gives you both offline queuing and immediate state sync without the four-way handshake cost. I also set the broker's max queued messages per client (for example, max_queued_messages 100 in Mosquitto) and implement back-pressure: if the queue fills, the broker should drop oldest or disconnect, not OOM.
Testing Failure Modes Before Deployment
I have learned to test three scenarios on the bench: 1) Kill power mid-QoS 2 publish and verify the handshake recovers after reboot. 2) Disconnect for longer than Session Expiry and confirm Session Present returns 0 and the device resubscribes correctly. 3) Subscribe after hours of retained publishes and measure the burst size and memory impact. The MQTT Specification is precise about these flows, but only a real disconnect test shows whether your client library correctly persists packet IDs to flash.
Finally, document your topic policy in one place: which topics are retained, which QoS each publisher and subscriber uses, and what Session Expiry each Client ID class gets. When that document exists, integrating a new microservice or replacing a broker becomes routine instead of risky.
Frequently Asked Questions
Does QoS guarantee end-to-end delivery from publisher to subscriber?
No. QoS is enforced hop-by-hop: publisher to broker, and broker to subscriber. A publisher sending at QoS 2 does not guarantee the subscriber receives at QoS 2 unless the subscriber also subscribed at QoS 2. The broker will downgrade delivery to the QoS of the subscription. For critical commands, ensure both sides request the same QoS.
Should I use retained messages for actuator commands?
Generally no. A retained command will be replayed to every new subscriber, including a device that just rebooted, causing unintended re-execution. Use retained messages for state and configuration, and use non-retained QoS 1 or QoS 2 for commands. If you must retain a command, include an expiry timestamp or nonce so the receiver can discard stale deliveries.
What is the difference between Clean Session in MQTT 3.1.1 and Clean Start + Session Expiry in MQTT 5?
In MQTT 3.1.1, Clean Session is a single boolean: 0 means persistent session forever until overwritten, 1 means no persistence. MQTT 5 splits this into Clean Start, which controls whether to resume an existing session at connect time, and Session Expiry Interval, which controls how long the broker keeps the session after disconnect. MQTT 5 allows timed expiry, for example keeping a session for one hour, which is much better for sleepy or decommissioned devices.
How do I prevent the broker from running out of memory for offline persistent sessions?
Set limits on the broker: maximum queued messages per client, maximum queued bytes, and message expiry interval. On the device, bound the number of in-flight QoS 1/2 messages and avoid subscribing at high QoS to high-throughput topics unless needed. For fleet management, use a short Session Expiry Interval for ephemeral clients and monitor broker metrics for queued message count per Client ID.