When I first migrated a line of aging packaging machines from OPC Classic to OPC UA, the motivation wasn't just to escape DCOM headaches — it was to gain a consistent way to describe what the machines actually *are*. OPC UA is often introduced as a protocol, but that framing misses the point. In my experience, its real value for industrial iot deployments is the combination of a platform-independent communication stack with a rich, object-oriented information model and built-in security. If you can model your process correctly and lock it down properly, the transport — whether Client/Server over TCP or PubSub over MQTT — becomes an implementation detail. This article covers how I approach opc ua information modeling and security for production systems, with lessons learned from deploying servers on both PLC-grade controllers and embedded Linux gateways.
Why OPC UA Displaced OPC Classic on the Shop Floor
OPC Classic relied on Microsoft COM/DCOM, which tied you to Windows, made remote networking fragile due to dynamic ports and firewall traversal, and offered no security beyond what DCOM provided. OPC UA, standardized as IEC 62541, was a clean-sheet redesign. It is service-oriented, platform-agnostic, and TCP-based with a binary encoding (UA Binary) for efficiency and an optional XML/JSON encoding for debugging and cloud integration.
In practice, this means you can run an OPC UA server on a Siemens S7-1500, a Beckhoff CX controller, an NXP i.MX8 gateway running embedded Linux, or even a Zephyr-based microcontroller acting as a field device gateway. The client doesn't care. I've had the same Node-RED and Python clients talk to all four without a single driver change, which was impossible in the OPC Classic world.
Service Sets That Actually Matter in Embedded Deployments
The OPC UA stack is large, but on constrained devices you only implement what you need. The profiles that matter most to me are:
Discovery: How clients find servers. On a flat OT network, Local Discovery Server (LDS) with mDNS works well. For larger sites, Global Discovery Server (GDS) with certificate management becomes essential — more on that in the security section.
Session Services: Creation of a SecureChannel and Session. This is where ApplicationInstanceCertificates are exchanged and SecurityPolicies are negotiated. If this fails, you'll see Bad_SecurityChecksFailed in your logs, which is almost always a certificate trust issue.
Subscription Services: MonitoredItems and Subscriptions are the core of efficient data exchange. Polling Read requests is wasteful. A properly configured Subscription with a publishing interval of 100-500 ms and a sensible queue size will reduce your network load by an order of magnitude compared to cyclic polling, especially when bridging to a Modern SCADA Architecture: From Legacy to Cloud-Connected Systems where bandwidth to the cloud is limited.
Method Services: Often overlooked, but critical if you want to expose actions, not just data. For example, exposing a CalibrateSensor or ExecuteMaintenanceCycle method with explicit input/output arguments is far cleaner than toggling a magic boolean variable.
Inside the OPC UA Address Space: Nodes, References, and Views
Every OPC UA server exposes an Address Space — a mesh of Nodes connected by References. Once you internalize this graph model, the rest of the standard makes sense. There are eight NodeClasses, but you will work with four daily:
Objects: Instances that give structure. Think CNC_Machine_12 or EnergyMeter. Objects don't hold values themselves; they organize.
Variables: The leaves that hold data. These have a Value, DataType, AccessLevel, and Historizing flag. A temperature sensor would be a Variable Node with DataType Double and EngineeringUnits.
Methods: Functions an Object can execute. They are called via the Call service.
ObjectTypes and VariableTypes: Type definitions that allow object-oriented modeling. This is where information modeling starts.
References are typed and directional. The most important are HasComponent, HasProperty, Organizes, and HasSubtype. In my experience, engineers new to OPC UA try to create deeply nested hierarchies using only HasComponent. I've found that using Organizes for folder-level grouping and HasComponent for true composition keeps browsing clients and Digital Twin Implementation: From Sensor Data to Simulation Models mappings much cleaner.
Attributes vs. Properties vs. Data Variables
Newcomers confuse Attributes, Properties, and Variables. Attributes are the fixed metadata defined by the NodeClass itself (e.g., NodeId, BrowseName, DisplayName, Value). Properties are a convention: they are Variables referenced with a HasProperty Reference and are meant to describe the parent Node (like SerialNumber or MaxTemperature). In the field, I keep Properties for static, descriptive metadata and use Data Variables for anything that changes at runtime or needs a timestamp and StatusCode. If your value needs historical logging or alarming, it must be a Data Variable, not a Property.
// Example: open62541 C - Creating a sensor object with a property and a data variable
UA_NodeId sensorNodeId = UA_NODEID_STRING(1, "Sensor_TI101");
UA_ObjectAttributes oAttr = UA_ObjectAttributes_default;
oAttr.displayName = UA_LOCALIZEDTEXT("en-US", "Temperature Sensor TI101");
UA_Server_addObjectNode(server, sensorNodeId,
UA_NODEID_NUMERIC(0, UA_NS0ID_OBJECTSFOLDER),
UA_NODEID_NUMERIC(0, UA_NS0ID_ORGANIZES),
UA_QUALIFIEDNAME(1, "TI101"),
UA_NODEID_NUMERIC(0, UA_NS0ID_BASEOBJECTTYPE),
oAttr, NULL, NULL);
// Property: SerialNumber (static)
UA_VariableAttributes pAttr = UA_VariableAttributes_default;
pAttr.displayName = UA_LOCALIZEDTEXT("en-US", "SerialNumber");
UA_String serial = UA_STRING("SN-2024-8841");
UA_Variant_setScalar(&pAttr.value, &serial, &UA_TYPES[UA_TYPES_STRING]);
pAttr.dataType = UA_TYPES[UA_TYPES_STRING].typeId;
UA_Server_addVariableNode(server, UA_NODEID_STRING(1, "TI101.SerialNumber"),
sensorNodeId, UA_NODEID_NUMERIC(0, UA_NS0ID_HASPROPERTY),
UA_QUALIFIEDNAME(1, "SerialNumber"),
UA_NODEID_NUMERIC(0, UA_NS0ID_PROPERTYTYPE), pAttr, NULL, NULL);
// Data Variable: Temperature (dynamic, historized)
UA_VariableAttributes vAttr = UA_VariableAttributes_default;
vAttr.displayName = UA_LOCALIZEDTEXT("en-US", "Temperature");
vAttr.dataType = UA_TYPES[UA_TYPES_DOUBLE].typeId;
vAttr.accessLevel = UA_ACCESSLEVELMASK_READ | UA_ACCESSLEVELMASK_HISTORYREAD;
UA_Double tempInit = 0.0;
UA_Variant_setScalar(&vAttr.value, &tempInit, &UA_TYPES[UA_TYPES_DOUBLE]);
UA_Server_addVariableNode(server, UA_NODEID_STRING(1, "TI101.Temperature"),
sensorNodeId, UA_NODEID_NUMERIC(0, UA_NS0ID_HASCOMPONENT),
UA_QUALIFIEDNAME(1, "Temperature"),
UA_NODEID_NUMERIC(0, UA_NS0ID_BASEDATAVARIABLETYPE), vAttr, NULL, NULL);
Notice the explicit use of HasProperty for the serial number and HasComponent for the live temperature. That distinction will be enforced by any OPC UA certification test tool and by companion specifications.
Crafting Robust Information Models Without Reinventing the Wheel
The biggest mistake I see is treating OPC UA like a bag of tags — mirroring every PLC address as a flat list of variables. That defeats the purpose. An effective information model adds semantics so a client knows that ColdEndTemperature is not just a Float, but a temperature with units, range, and a relationship to a specific extruder zone.
OPC UA provides a base information model in Namespace 0, and then Companion Specifications extend it for domains: OPC 40001 for Machinery, OPC 40083 for Pumps, PackML (OMAC) for packaging machines, AutoID for RFID/Scanners, and ISA-95 style models. Before defining a single custom Node, check the OPC Foundation's Companion Specification list. In my experience, 70% of what you need is already standardized. Starting from PackML for a fill-and-seal line saved us three weeks of modeling and made our MES integration plug-and-play.
Namespace Strategy and NodeId Design
Keep your Namespace at index 1 or higher; Namespace 0 is reserved. For NodeIds, I've settled on String NodeIds for most industrial assets because they survive reboots and code regeneration. Numeric NodeIds are faster on constrained stacks but are fragile if auto-generated.
Define your types first, then your instances. For example, define a ExtruderZoneType (ObjectType) with components ActualTemperature, Setpoint, and HeaterStatus, then instantiate Zone_1 through Zone_5 from it. This guarantees consistency and allows a client to discover all zones via a simple Browse for Objects of that type.
<!-- Fragment of a NodeSet2.xml Information Model -->
<UAObjectType NodeId="ns=1;i=1001" BrowseName="1:ExtruderZoneType">
<DisplayName>ExtruderZoneType</DisplayName>
<References>
<Reference ReferenceType="HasSubtype" IsForward="false">i=58</Reference> <!-- BaseObjectType -->
</References>
</UAObjectType>
<UAVariable NodeId="ns=1;i=6001" BrowseName="1:ActualTemperature"
ParentNodeId="ns=1;i=1001" DataType="Double">
<DisplayName>ActualTemperature</DisplayName>
<References>
<Reference ReferenceType="HasComponent" IsForward="false">ns=1;i=1001</Reference>
<Reference ReferenceType="HasTypeDefinition">i=63</Reference> <!-- BaseDataVariableType -->
</References>
</UAVariable>
<UAObject NodeId="ns=1;s=Extruder.Zone1" BrowseName="1:Zone1">
<DisplayName>Zone1</DisplayName>
<References>
<Reference ReferenceType="HasTypeDefinition">ns=1;i=1001</Reference>
<Reference ReferenceType="HasComponent" IsForward="false">ns=1;s=Extruder</Reference>
</References>
</UAObject>
Tools like UA Modeler and open-source nodeset compilers can generate C or Python code from this XML. When running on a resource-constrained device, consult the FreeRTOS Documentation for memory management patterns if your OPC UA stack shares the heap with real-time tasks — I've seen poorly sized nodesets consume 400KB of RAM before a single connection is made.
Practical Rules for Model Longevity
One: Version your Namespace URI (e.g., http://yourcompany.com/Extruder/1.1) and never change the semantics of an existing NodeId. Add new Nodes instead. Two: Use ModellingRules like Mandatory and Optional to make type compliance explicit. Three: Add Description attributes and EngineeringUnits (EUInformation) to every analogue variable. The MES engineer integrating your server six months later will thank you.
Mapping Communication Patterns to Industrial IoT Architectures
OPC UA defines two distinct communication paradigms, and choosing the right one determines your architecture.
Client/Server: Classic request-response over OPC UA TCP (opc.tcp://). The client creates a Session, Subscribes to data, and the server publishes Notifications. This is perfect for SCADA, HMI, and MES — any system that needs reliable, contextual access with browsing, history, and method calls. Latency is typically 10-50 ms, which is fine for supervisory control but not for closed-loop control.
PubSub: Introduced in OPC UA 1.04, PubSub decouples publishers and subscribers via a MessageOriented Middleware. The publisher encodes DataSetMessages and sends them via UDP multicast (for local real-time), MQTT, or AMQP. This is ideal for cloud ingestion and edge analytics where many consumers need the same telemetry without creating thousands of Client/Server sessions. According to the MQTT Specification, MQTT 5.0 provides the underlying acknowledgement and retained message semantics that OPC UA PubSub leverages when you select MQTT as the transport.
In my experience, the hybrid approach works best: use Client/Server within the cell for configuration, diagnostics, and method execution, and PubSub over MQTT for streaming telemetry to an edge aggregator or cloud broker. This mirrors how Industrial Ethernet: PROFINET, EtherCAT and TSN for Real-Time Control separates the hard real-time fieldbus from the information layer — OPC UA is not a replacement for PROFINET IRT or EtherCAT, it is the northbound aggregation layer above them.
# Python asyncua - Efficient subscription vs. inefficient polling
import asyncio
from asyncua import Client
async def main():
async with Client("opc.tcp://192.168.1.50:4840") as client:
# Bad: Polling every 500ms
# while True:
# val = await client.get_node("ns=1;s=Extruder.Zone1.ActualTemperature").read_value()
# await asyncio.sleep(0.5)
# Good: Subscription with 250ms publishing interval
node = client.get_node("ns=1;s=Extruder.Zone1.ActualTemperature")
sub = await client.create_subscription(250, handler)
handle = await sub.subscribe_data_change(node)
await asyncio.sleep(300) # run for 5 minutes
await sub.unsubscribe(handle)
class handler:
async def datachange_notification(self, node, val, data):
print(f"{node} -> {val} @ {data.monitored_item.Value.SourceTimestamp}")
if __name__ == "__main__":
asyncio.run(main())
Note the publishing interval is a hint to the server; the server may revise it. Always check the revised value returned by CreateSubscription. On embedded servers like open62541 or node-opcua, I set a minimum publishing interval of 100 ms to prevent a misbehaving client from overwhelming the controller.
OPC UA Security That Survives a Factory Audit
OPC UA security is comprehensive but not automatic. Out of the box, many sample servers run with SecurityPolicy None and no certificate validation — which will fail any credible security audit. A production deployment needs to get four layers right.
1. Transport Security and SecurityPolicies: Every Endpoint advertises a SecurityPolicy and SecurityMode. SecurityMode None means no signing or encryption. For production, I disable it entirely. The viable policies are:
| SecurityPolicy | Encryption & Signing | Certificate Requirement | Typical Use Case |
|---|---|---|---|
| None | None | No certificate | Lab/commissioning only |
| Basic256Sha256 | AES-256 + RSA-OAEP + SHA-256 | ApplicationInstanceCertificate required | Current recommended default |
| Aes128_Sha256_RsaOaep | AES-128 + RSA-OAEP + SHA-256 | ApplicationInstanceCertificate required | New installations, slightly less CPU load |
| Aes256_Sha256_RsaPss | AES-256 + RSA-PSS + SHA-256 | ApplicationInstanceCertificate required | Highest security, newer stacks only |
On constrained hardware, Aes128_Sha256_RsaOaep offers a good balance. I've measured a 15-20% reduction in SecureChannel handshake time on a Cortex-A53 compared to Basic256Sha256, which matters if devices reconnect frequently.
2. Application Authentication (X.509 Certificates): OPC UA uses X.509 v3 certificates to authenticate applications, not just users. Each server and client has an ApplicationInstanceCertificate with a private key. The trust decision happens via a TrustList — a folder of trusted peer certificates and CRLs. In my experience, manually copying .der files via USB stick works for 5 nodes but collapses at 50. For fleet scale, implement GDS (Global Discovery Server) with Push certificate management or integrate with your existing PKI. Certificates should be at least 2048-bit RSA (4096-bit preferred) with a validity of 1-2 years and proper SubjectAltName containing the ApplicationUri and hostnames/IPs.
3. User Authentication and Authorization: Even with a trusted Application, you still need UserIdentityTokens. OPC UA supports Anonymous (disable it), UserName/Password (hashed with the SecureChannel encryption), X509 Certificate, and IssuedToken (Kerberos/JWT). After authentication, authorization is enforced via Role-Based Access Control (introduced in 1.05) and per-Node AccessRestrictions and RolePermissions. For example, an Operator role may have Read and Call rights on Method Nodes, but only Maintenance may Write to Setpoints.
4. Hardening the Embedded Stack: Beyond OPC UA itself, disable anonymous access, set minimum SecurityPolicy, limit Session count (e.g., max 10 sessions on a gateway), and implement AuditEvents for login failures and Write operations. I've found that logging Security Audit Events to a local ring buffer and forwarding them via Syslog catches brute-force attempts early. If you are building the server on Zephyr or FreeRTOS, refer to the Zephyr Project Documentation for secure storage of private keys in the trusted execution environment or secure element — storing a .pem file on an unauthenticated flash filesystem is not secure.
Diagnosing OPC UA Throughput Limits on Embedded Gateways
The last deployment that bit me was a gateway aggregating 800 variables from three PLCs and serving them to four MES clients and a cloud publisher. The CPU wasn't the bottleneck — the Subscription handling was.
MonitoredItem Queue Sizing: Each MonitoredItem has a queue (default 1 or 10). If your sampling interval is faster than your publishing interval, values will queue or be discarded based on DiscardOldest. A queue of 10 with a 100 ms sampling and 1 s publishing interval can burst 10 notifications per publish. Multiply by 800 items and you flood the TCP window. I tune queues to 1 for most telemetry with StatusCode-driven triggers and use a queue of 5-10 only for critical alarms where every edge matters.
Deadbands and Filters: Use DataChangeFilters. An Absolute Deadband of 0.5 degrees on a noisy temperature signal can cut notification traffic by 60% without losing process relevance. For counters or energy meters, use PercentDeadband.
Max Nodes Per Read/Browse: Embedded servers often cap operations like MaxNodesPerRead at 100-500 to prevent memory exhaustion from a single giant Read request. Well-behaved clients should split bulk reads. Wireshark with the OPC UA dissector is your best friend here — filter by opcua and watch for Bad_TooManyOperations responses.
Memory Footprint: A minimal open62541 server with 2000 Nodes and two Subscriptions needs roughly 1.2 MB Flash and 400-600 KB RAM. Add historical access or PubSub and that doubles. On an STM32H7 with 480 KB RAM, I had to disable Historian and use dynamic Node creation only for the active production recipe. Design your information model to be lazy-loaded if you are memory-constrained.
Finally, test worst-case client behavior, not just the happy path. I've used a Python script that opens 20 parallel Sessions and creates 1000 MonitoredItems each — exactly what a misconfigured SCADA driver will do. If your gateway survives that without watchdog resets, it will survive the plant.
Frequently Asked Questions
Can OPC UA replace PROFINET or EtherCAT for real-time machine control?
No. OPC UA is not a fieldbus for hard real-time, deterministic I/O. PROFINET IRT, EtherCAT, and TSN-based networks provide cycle times below 1 ms with jitter in microseconds for motion control. OPC UA excels at the information layer above — supervision, configuration, diagnostics, and cloud integration — with typical latencies of 10-100 ms for Client/Server and 1-10 ms for PubSub over TSN/UDP multicast. I always keep the real-time control loop on the fieldbus and expose its status and parameters via OPC UA.
How do I choose between String and Numeric NodeIds for my information model?
Use String NodeIds (e.g., ns=1;s=Pump101.Pressure) for human-readable, stable identifiers that survive code regeneration and are easy to map to digital twins. Use Numeric NodeIds (e.g., ns=1;i=5001) only when you need maximum performance on highly constrained stacks or when auto-generating large arrays of identical nodes where lookup speed matters. In my experience, String NodeIds reduce commissioning errors and make logging and diagnostics far clearer, with negligible performance impact on modern ARM controllers.
What is the most common cause of Bad_SecurityChecksFailed when connecting a client?
Almost always a certificate trust issue. The client’s ApplicationInstanceCertificate is not in the server’s trusted store, or vice versa, or the ApplicationUri in the certificate does not match the ApplicationUri configured in the endpoint. Check that the certificate’s SubjectAltName includes the correct hostname/IP, that the current time is synchronized (certificates appear invalid if the device clock is wrong), and that you have copied the certificate to the trusted folder — not the rejected folder. Using GDS for automated trustlist management eliminates most of these manual errors at scale.
Should I use OPC UA PubSub over MQTT or a direct Client/Server subscription for cloud telemetry?
For telemetry to the cloud, I prefer OPC UA PubSub over MQTT. It preserves the OPC UA DataSet metadata (DataTypes, timestamps, StatusCodes) and allows efficient binary encoding while leveraging MQTT’s broker scalability and retained messages. Direct Client/Server subscriptions require a persistent TCP session per consumer, which does not scale to hundreds of edge consumers. Reserve Client/Server for on-premise SCADA and for browsing, historical reads, and method calls that PubSub does not support.