Authenticating an IoT device is not the same problem as authenticating a user. A device has no password to reset, no human to complete a second factor prompt, and often no secure channel to recover once its identity is compromised. In production fleets ranging from a few hundred industrial gateways to several million battery-powered sensors, I have seen authentication shortcut more often than any other security control, usually in the name of time-to-market. The result is predictable: cloned devices, spoofed telemetry, and backend systems that cannot trust their own data. This article covers how modern embedded fleets establish verifiable identity using X.509 certificates and pre-shared keys (PSK), where each approach breaks down, and how to implement provisioning, rotation, and validation so that identity survives for the full device lifecycle.
Why Mutual TLS with X.509 Remains the Baseline for Fleet Identity
In my experience, any fleet expected to operate for more than three years should start with mutual TLS (mTLS) as the default assumption, even if you ultimately decide to deploy PSK on a subset of endpoints. The reason is not cryptography fashion. An X.509 certificate binds a public key to a device identity with attributes that the server can verify without sharing a secret. The private key never leaves the device, and the server validates the chain to a trusted root. This gives you three properties PSKs cannot: non-exportable identity, scalable revocation, and delegation through intermediate CAs.
The practical distinction shows up during manufacturing and ownership transfer. With certificates, a contract manufacturer can inject a device certificate signed by your intermediate CA without ever seeing your root key. With PSK, someone has to generate, distribute, and store a symmetric secret on both ends, and that secret must be protected in transit. For a fleet of 10,000 devices, that is manageable. For 500,000, key distribution becomes a logistics and HSM problem that is often more expensive than a PKI.
mTLS also maps cleanly to how cloud IoT brokers authenticate. AWS IoT Core, Azure IoT Hub, and most private vernemq or EMQX deployments expect the device to present a client certificate during the TLS handshake. The broker extracts the Common Name or Subject Alternative Name and uses it as the logical device ID for policy evaluation. This binding between the TLS layer and the application layer eliminates a whole class of confused-deputy bugs where a device authenticates as one identity but publishes as another.
Anatomy of a Device Certificate That Actually Scales
A device certificate should be an automated artifact, not a hand-crafted file. I keep the subject deliberately minimal: CN=device-serial-001A3F plus a SAN that encodes the immutable device ID. Extended attributes like product line, hardware revision, and manufacturing site belong in custom X.509v3 extensions or in a device registry, not overloaded into the OU field. Keep validity periods short, 1-2 years maximum for operational certificates, with a longer-lived factory bootstrap certificate used only for initial enrollment.
Key generation must happen on-device or inside a secure element. If you generate keys on a workstation and copy them over USB during production, you have already created an extractable secret that can be cloned. Modern parts like the ATECC608, ST STSAFE, or NXP SE050 will generate a P-256 or RSA-2048 keypair internally and never expose the private key. The device then outputs a Certificate Signing Request (CSR). Your signing service returns a certificate without the private key ever existing outside silicon.
# Generate a P-256 keypair and CSR on a Linux gateway (or inside a secure element via PKCS#11)
openssl ecparam -name prime256v1 -genkey -noout -out device.key
openssl req -new -key device.key -out device.csr \
-subj "/CN=sensor-4f8a9c2e/O=Acme Sensors/C=US" \
-addext "subjectAltName=URI:urn:acme:device:4f8a9c2e"
# Sign with your intermediate CA (valid for 397 days)
openssl x509 -req -in device.csr -CA intermediate.crt -CAkey intermediate.key \
-CAcreateserial -out device.crt -days 397 -sha256 \
-extfile <(printf "subjectAltName=URI:urn:acme:device:4f8a9c2e\nextendedKeyUsage=clientAuth")
On the verification side, never disable hostname verification or set your TLS library to VERIFY_NONE because your test broker used a self-signed cert. Flash the CA bundle into read-only storage and pin your intermediate if you control the PKI. The MQTT Specification itself does not define authentication, it relies entirely on the underlying TLS session, which means a misconfigured TLS stack silently undermines your entire session management.
Provisioning X.509 Credentials: HSMs, Secure Elements and Factory Floors
Provisioning is where most X.509 deployments fail. I have audited lines where certificates were correctly architected on a whiteboard but then injected via an unencrypted Python script over UART on the factory floor, with operator workstations storing private keys on desktop folders. The threat model for provisioning must assume the contract manufacturer is untrusted.
There are two patterns that consistently work. First is factory injection with an offline HSM. The HSM holds the intermediate signing key and is kept in a locked enclosure that only exposes a CSR-in, certificate-out interface. Operators submit CSRs generated by the secure element and receive signed certificates. No one on the line handles keys. Second is Just-In-Time Provisioning (JITP) or Just-In-Time Registration (JITR) where the device ships with a bootstrap certificate signed by a factory CA, and on first connection the cloud verifies that bootstrap chain and automatically registers the device, often issuing an operational certificate via EST or SCEP.
JITP scales better for high-volume consumer devices because you do not need to pre-register every serial number before manufacturing. AWS IoT's template-based provisioning and Azure Device Provisioning Service both support this, but you must enforce that the bootstrap certificate can only call the provisioning endpoint and nothing else. I've found that allowing a factory certificate to publish telemetry creates a permanent backdoor if that factory CA is ever compromised.
Implementing Secure Storage on Zephyr and FreeRTOS
On constrained RTOS platforms, the TLS stack integration dictates how you store credentials. With Zephyr Project Documentation you typically use the PSA Crypto or TLS credential manager, while on FreeRTOS you use the PKCS#11 abstraction over your secure element. Do not store device.key as a plaintext file in LittleFS. Use the key storage API so the key handle, not the key material, is passed to mbedTLS.
/* Zephyr / mbedTLS: loading credentials from the TLS credential subsystem */
const unsigned char ca_cert[] = {
#include "ca_chain_der.inc"
};
tls_credential_add(CA_CERTIFICATE, TLS_CREDENTIAL_CA_CERTIFICATE,
ca_cert, sizeof(ca_cert));
/* Private key is referenced by ID, not buffered in RAM */
tls_credential_add(TLS_CREDENTIAL_PRIVATE_KEY,
SECURE_ELEMENT_KEY_ID, NULL, 0);
/* FreeRTOS + coreMQTT: PKCS#11 private key label */
CK_OBJECT_HANDLE xPrivKey;
CK_ATTRIBUTE xTemplate = { CKA_LABEL, "Device PrivKey", 15 };
C_FindObjectsInit(xSession, &xTemplate, 1);
C_FindObjects(xSession, &xPrivKey, 1, &ulCount);
In load tests, I have measured handshake memory overhead of 28-45 kB for ECDSA certificates versus 60-90 kB for RSA-2048 on Cortex-M4 platforms. If you are within 20 kB of your RAM limit, that difference decides whether mTLS is feasible without external PSRAM.
When Pre-Shared Keys Make Pragmatic Sense on Constrained Hardware
PSK is often dismissed as insecure by default. I disagree. A properly implemented TLS-PSK or DTLS-PSK with a 256-bit random key per device, stored in a secure element and provisioned over an encrypted channel, is far stronger than a shared certificate used by every device. The issue is not the cryptography, it is key management at scale.
Where PSK excels is in ultra-constrained, battery-powered endpoints that sleep for hours and wake to send 20 bytes over DTLS. The handshake for TLS-PSK with AES-128-CCM-8 is 3-4 times smaller and 40% faster than ECDHE-ECDSA, and it eliminates certificate parsing, ASN.1 handling, and chain validation which can be heavy on a Cortex-M0+. This is a primary reason NB-IoT water meters and some LoRaWAN Network Architecture: Gateways, Network Server and Join Procedures edge bridges use PSK-derived session keys for their backhaul.
My rule is: use PSK only when you meet three conditions. One, each device gets a unique, high-entropy PSK generated by an HSM, never derived from a serial number. Two, the PSK identity hint maps one-to-one to your device registry and you enforce that mapping on the server. Three, you have an automated rotation plan that does not require a truck roll. If you cannot meet all three, use certificates.
Protocol choice influences this decision heavily. If you are weighing transport overhead and retransmission behavior, CoAP vs MQTT: Choosing the Right IoT Protocol for Constrained Devices provides a useful breakdown of how DTLS-PSK with CoAP compares to TLS certificates with MQTT on lossy links. I have typically found that CoAP + DTLS-PSK is the only viable option on 802.15.4 meshes where MQTT over TLS would keep the radio active too long.
Hardening TLS-PSK Against Identity Spoofing
The classic PSK failure is using one fleet-wide key with an identity string like psk-identity: factory-default. Anyone who extracts one device recovers the fleet. Always use the PSK identity as a lookup key, not as a secret, and verify it on the server against an allow-list. On the device side, store the PSK using the same secure element interface you would for a private key. The Zephyr TLS stack and mbedTLS both support MBEDTLS_KEY_EXCHANGE_PSK_ENABLED without needing X.509 parsers at all, which saves flash.
Implementing Certificate Rotation Without Bricking Field Devices
A static certificate that never rotates is a ticking incident. Private keys leak, algorithms age, and CA hierarchies change. I've found that designing rotation on day one is 10x cheaper than retrofitting it after your first SHA-1 deprecation or CA compromise.
Implement a dual-certificate slot. Slot A holds the active operational certificate, Slot B is staging. Rotation happens in three phases: Provision new certificate to Slot B while keeping Slot A active, attempt TLS connection with Slot B and fall back to Slot A if the broker rejects it, and only erase Slot A after observing a successful connection window. This atomic approach protects against power loss during flash writes, which is the most common cause of devices losing identity and never reconnecting.
Automate issuance via EST (Enrollment over Secure Transport, RFC 7030) or a custom HTTPS enrollment endpoint authenticated with the current mTLS session. The device authenticates with its existing certificate to request a new one, proving possession of the old private key. The server validates that the device is still owned by the same tenant before signing.
/* Simplified EST re-enrollment flow on device */
est_ctx_t *ctx = est_client_init(ca_chain, operational_cert, priv_key_handle);
if (est_enroll(ctx, "/.well-known/est/simplereenroll", &new_csr, &new_cert) == 0) {
/* Write to inactive slot first */
nvs_write(NVS_SLOT_B_CERT, new_cert, cert_len);
if (tls_probe_with_cert(NVS_SLOT_B_CERT) == TLS_OK) {
nvs_erase(NVS_SLOT_A_CERT);
nvs_copy(NVS_SLOT_B_CERT, NVS_SLOT_A_CERT);
reboot_with_new_identity();
} else {
/* Keep existing identity, report failure via shadow */
report_rotation_failed(new_cert.serial);
}
}
Plan for clock skew. Embedded devices without an RTC will boot at epoch 0 and fail certificate validation because notBefore is in the future. Either sync time via NTP or authenticated CoAP time before validating, or issue bootstrap certificates with a notBefore anchored at manufacturing date. The FreeRTOS Documentation recommends verifying time sync before the first TLS handshake for exactly this reason.
Binding Device Identity to Protocol Sessions: MQTT, CoAP and LwM2M Realities
A valid TLS handshake is only half of authentication. You must bind the verified identity to the application session and enforce it. On MQTT, this means the broker's authorizer must check that the certificate's SAN matches the clientId and the topics the device is allowed to publish to. I enforce a topic policy like devices/{deviceId}/telemetry where {deviceId} must equal the certificate identity. Without this, a compromised device can publish telemetry as any other device by changing its clientId.
For a deeper look at how sessions, clean start flags, and retained messages interact with authentication state, see MQTT Protocol Deep Dive: QoS Levels, Retained Messages and Session Management. A subtle bug I have seen repeatedly is that a device reconnects with a new certificate but resumes an old persistent session that was authorized under the old identity, bypassing new policy.
With CoAP over DTLS, identity binding is less standardized. Use the OSCORE or DTLS connection ID extension to maintain session context across sleepy device wakeups, and map the DTLS PSK identity or certificate fingerprint to your resource ACL. LwM2M 1.1 formalizes this with Security and Server objects that can hold either PSK, RPK or X.509 credentials and negotiate the mode during bootstrap. If you support multiple modes, do not allow downgrade. A device provisioned for X.509 must not be permitted to connect via PSK with a guessed identity.
Handling Identity in BLE and Local Provisioning
Local onboarding often happens over BLE before the device has Wi-Fi credentials. Use BLE secure pairing with LE Secure Connections and then transfer Wi-Fi and certificate provisioning data over an encrypted GATT characteristic, not over open advertising. The design patterns for GATT service isolation discussed in BLE Development: GATT Services, Advertising and Connection Management apply directly here. After provisioning, wipe the BLE provisioning service so it cannot be re-used to extract credentials.
Validating and Revoking Identity: OCSP Stapling, CRLs and Cloud-Side Policy
Validation is not only about chains. You need to answer: is this certificate still trusted right now? Traditional CRLs and OCSP do not work well on constrained devices that may be offline. Devices cannot reliably fetch a multi-megabyte CRL, and OCSP requires an extra round trip that kills battery life.
Therefore, enforce revocation primarily on the server. Keep the device validation path simple: validate chain, expiry, and key usage. Let the broker query a real-time device registry or policy engine during the handshake. If a device is revoked, the broker closes the TLS session with bad_certificate or access_denied. For high-security installations, enable OCSP stapling where the server provides a time-stamped OCSP response during the handshake, so the device can verify server identity without an extra lookup.
Short-lived certificates are a more practical revocation strategy for IoT than CRL distribution. If operational certificates live for 7-30 days and are auto-renewed, revocation is simply non-renewal. An attacker who compromises a key has a limited window before the certificate expires and cannot be renewed without the provisioning service authorizing it.
| Method | Secret Type | Provisioning Complexity | Scalability to 1M+ Devices | Best Fit |
|---|---|---|---|---|
| X.509 mTLS | Asymmetric private key (non-exportable) | High initial (PKI, HSM), low per-device | Excellent with automated CA | Gateways, industrial controllers, long-life devices |
| TLS-PSK / DTLS-PSK | 256-bit symmetric key | Low initial, high per-device logistics | Difficult without automated key vault | Battery sensors, NB-IoT, 802.15.4 meshes |
| Raw Public Key (RPK) | Asymmetric private key, no certificate | Medium (key pinning required) | Moderate, manual trust anchor mgmt | Closed systems, Lightweight bootstrapping |
| Token / JWT (MQTT password) | Bearer token derived from key | Low | Good if token service is scalable | Cloud-brokered devices with frequent re-auth |
Finally, log and alert on authentication anomalies at the fleet level. Repeated failed handshakes from one region, a spike in unknown certificate serials, or a device suddenly presenting a different CA chain are early indicators of cloning or factory leakage. I feed TLS handshake failures into the same pipeline as telemetry, so an operational dashboard shows identity health alongside sensor data. Identity without observability is just cryptography you hope is working.
Frequently Asked Questions
Can I use a single wildcard certificate for all devices to simplify provisioning?
No, and this is the most common antipattern I remediate. A shared wildcard private key means extracting one device compromises the entire fleet and you cannot revoke a single device without rotating every device. Always issue a unique certificate or PSK per device with a unique private key generated inside a secure element. The operational overhead is higher, but it is the only design that contains a breach.
How do I handle X.509 on a device with only 64 kB RAM?
Use ECDSA P-256 instead of RSA, strip your CA bundle to only the intermediates you trust, and use a TLS library configured for minimal handshake buffers. MbedTLS with MBEDTLS_ECP_DP_SECP256R1_ENABLED and session resumption can fit in under 35 kB. If that still does not fit, switch to DTLS-PSK with AES-CCM and push certificate validation to the gateway that aggregates the constrained nodes.
What is the safest way to rotate a compromised intermediate CA?
Do not revoke the intermediate abruptly. Build a new intermediate, start issuing new device certificates from it, and deploy a broker trust store that trusts both old and new intermediates concurrently. Roll devices through the dual-slot rotation described above. Once your telemetry confirms that 99.5% of active devices have migrated, remove the compromised intermediate from the trust store and revoke any remaining certificates. This avoids mass disconnection.
Should I use self-signed device certificates with manual pinning?
Self-signed certificates eliminate CA management but replace it with manual pinning distribution, which is fragile. You must securely distribute every device's self-signed certificate or public key to the server before it can connect, and any re-key requires a server update. In my experience, self-signed is only viable for lab prototypes or closed fleets managed over a private APN. For production, a lightweight internal CA with automated issuance is more maintainable.