CAN Bus Protocol: Automotive and Industrial Communication Networks

When a production line halts because a single sensor node floods the network with error frames, or when an automotive ECU stops responding at -30°C despite passing bench tests, you quickly learn that CAN is not just another serial protocol you can treat like UART with a different connector. I have spent the last decade bringing CAN-based sensor networks into service on factory floors, agricultural machines, and battery management systems, and the difference between a bus that runs for years without intervention and one that requires constant resets almost always comes down to details that datasheets gloss over: termination and stub length discipline, bit timing margins, identifier allocation strategy, and how your firmware handles error confinement. This article shares that practical layer — the decisions, measurements, and code patterns that make CAN reliable at scale for automotive and industrial communication.

CAN Physical Layer Realities: Dominant Bits, Termination and Transceiver Selection

In my experience, most CAN reliability problems are physical layer problems misdiagnosed as software bugs. CAN uses differential signaling on CAN_H and CAN_L with a nominal 120-ohm characteristic impedance. A logical '0' (dominant) drives the pair apart — CAN_H toward ~3.5V and CAN_L toward ~1.5V — while a logical '1' (recessive) lets both lines float to ~2.5V. Because dominant overwrites recessive, the bus implements a wired-AND that arbitration depends on. If your transceiver cannot drive that differential voltage under load, arbitration fails in subtle ways.

Two details I now enforce on every hardware review: proper termination and stub management. A correctly terminated CAN bus has a 120-ohm resistor at each physical end, yielding ~60 ohms measured between CAN_H and CAN_L with power off. I have seen installations with a single 120-ohm resistor, or three resistors, or 120-ohm resistors placed in the middle of the bus. All of them passed at room temperature on a short bench harness and failed on a 40-meter run with variable loads. The fix is simple but non-negotiable: place termination at the two extreme ends only, and measure the resistance before first power-on.

Transceiver Choice and Isolation Trade-offs

For 3.3V embedded nodes, parts like the TCAN332, MCP2562, or ISO1050 (isolated) are common choices. I've found that the standby/silent mode pins cause more field issues than they solve when misused. If you tie STB to ground to force normal mode, you lose the ability to keep a faulty node from jamming the bus. If you control it from the MCU, make sure your bootloader drives it to a safe state during startup — I have debugged a system where the bootloader left the transceiver in dominant state for 800ms while waiting for flash initialization, triggering bus-off on every other node.

In industrial environments with long runs, ground potential differences between cabinets can exceed several volts. A non-isolated transceiver will still communicate, but common-mode transients from motor drives can exceed its ratings. For any run that leaves a cabinet or shares a ground with power electronics, I specify galvanic isolation (either an isolated transceiver or separate digital isolator + transceiver) and isolated power. The additional cost is far less than a day of downtime tracing intermittent error frames that only appear when a VFD starts. The Zephyr Project Documentation has excellent reference schematics for isolated CAN interfaces that match what I use in production.

Stub Length and Topology Discipline

CAN is specified as a linear bus, not a star. Each drop (stub) from the main trunk should be as short as possible — under 0.3m for 1 Mbps, and under 1m for 500 kbps and lower is a safe rule. I've audited agricultural vehicles where installers created a 4-meter stub to reach a sensor on an implement. At 250 kbps it worked intermittently, but error counters incremented steadily. The solution was to re-route the backbone through the sensor and keep the stub under 30cm, eliminating reflections without changing baud rate.

Message Arbitration and Identifier Strategy Beyond the Textbook Example

Every CAN frame starts with an identifier that serves two purposes: bus arbitration and receiver filtering. During arbitration, nodes transmit their identifier bit-by-bit, monitoring the bus. If a node transmits a recessive bit but reads dominant, it has lost arbitration and stops transmitting. This means numerically lower identifiers have higher priority and will always win. That mechanism is elegant and deterministic, but it can starve lower-priority traffic if you allocate identifiers poorly.

In one building automation retrofit, the integrator assigned identifiers sequentially by installation order: the first damper controller got 0x100, the last got 0x7FF. Under light load everything worked. Under heavy load — when a fire alarm event triggered dozens of sensors simultaneously — the nodes with high identifiers experienced latencies exceeding 200ms because they repeatedly lost arbitration. The control loop, tuned for 50ms updates, became unstable. We redesigned the identifier map around function criticality, not installation order.

Designing an Identifier Map for Sensor Networks

For sensor interfacing, I recommend partitioning your 11-bit (or 29-bit) space with intent. For example, using 11-bit standard frames, reserve 0x000-0x0FF for network management and emergency frames (bus load shedding, heartbeat errors), 0x100-0x2FF for high-priority periodic sensor data (safety interlocks, current/voltage measurements), 0x300-0x5FF for normal periodic data (temperature, pressure, environmental sensors), and 0x600-0x7FF for configuration, diagnostics, and non-time-critical transfers. If you expect more than a few dozen nodes, move to extended 29-bit identifiers and structure them as: priority (3 bits) | source address (8 bits) | message type (10 bits) | destination (8 bits), similar to J1939.

Keep in mind that filtering happens in hardware. Most CAN controllers provide acceptance filters that can ignore irrelevant identifiers before they interrupt the CPU. On an STM32 or NXP S32K, you get 14-28 configurable filter banks. Use them. I have seen firmware that accepted every frame and filtered in software, burning 30% CPU at moderate bus loads. That approach collapses at bus loads above 50%. A well-configured hardware filter set can reduce interrupt load by an order of magnitude.

// STM32 HAL: Configure hardware filter to accept only sensor data 0x300-0x31F
// and emergency frames 0x001-0x00F, reject everything else in hardware
CAN_FilterTypeDef sFilterConfig;

sFilterConfig.FilterBank = 0;
sFilterConfig.FilterMode = CAN_FILTERMODE_IDMASK;
sFilterConfig.FilterScale = CAN_FILTERSCALE_32BIT;
sFilterConfig.FilterIdHigh = (0x300 << 5);      // ID 0x300
sFilterConfig.FilterIdLow = 0x0000;
sFilterConfig.FilterMaskIdHigh = (0x7E0 << 5);  // Mask for 0x7E0 -> accepts 0x300-0x31F
sFilterConfig.FilterMaskIdLow = 0x0000;
sFilterConfig.FilterFIFOAssignment = CAN_RX_FIFO0;
sFilterConfig.FilterActivation = ENABLE;
HAL_CAN_ConfigFilter(&hcan1, &sFilterConfig);

sFilterConfig.FilterBank = 1;
sFilterConfig.FilterIdHigh = (0x001 << 5);
sFilterConfig.FilterMaskIdHigh = (0x7F0 << 5);  // Accepts 0x000-0x00F
HAL_CAN_ConfigFilter(&hcan1, &sFilterConfig);

Unlike point-to-point protocols, CAN is multi-master by nature. This is fundamentally different from I2C Protocol Tutorial: Addressing, Clock Stretching and Multi-Master where multi-master exists but is rarely used and arbitration is more of an exception. On CAN, you should design as if multiple nodes will always contend. If your sensor network also includes SPI-attached high-speed sensors feeding a CAN gateway, review SPI Communication: Full-Duplex Data Transfer for High-Speed Sensors to keep sampling jitter from propagating onto the CAN schedule.

Bit Timing Configuration: Why Your 500 kbps Bus Fails at Temperature Extremes

Bit timing is the most common source of temperature-dependent CAN failures, and the least understood. A CAN bit is divided into segments: SYNC_SEG (1 TQ), PROP_SEG, PHASE_SEG1 and PHASE_SEG2, sampled at a point between PHASE_SEG1 and PHASE_SEG2. The bit time is the sum of these segments measured in time quanta (TQ), where TQ = prescaler / CAN_CLK. Two settings dominate robustness: sample point and synchronization jump width (SJW).

Textbook configurations often set the sample point at ~75% for classical CAN. I've found that for industrial buses with long cables and opto-isolation adding propagation delay, moving the sample point to 80-87.5% improves margin, provided your oscillator tolerance supports it. The critical constraint is propagation delay: the time from a transmitter driving dominant to that dominant being seen by every other node must be less than PROP_SEG.

Calculating Robust Timing for 500 kbps on a 80 MHz Clock

Assume CAN clock = 80 MHz, desired bit rate = 500 kbps. Choose 16 TQ per bit for granularity. Then TQ duration = 2 * prescaler / 80MHz; for 16 TQ per bit, nominal bit time = 2.0us (1/500k). Use prescaler = 5: TQ = 62.5ns, Bit Time = 16 * 62.5ns = 1.0us? That's wrong — let's calculate precisely: With prescaler = 10, CAN_CLK_TQ = 80MHz/10 = 8MHz, TQ = 125ns, 16 TQ = 2.0us = 500 kbps exactly. Now allocate: SYNC_SEG = 1 TQ, PROP_SEG = 6 TQ, PHASE_SEG1 = 5 TQ, PHASE_SEG2 = 4 TQ. Sample point = (1+6+5)/16 = 75%. SJW = 2-3 TQ minimum to track temperature drift. In my experience, sharing this calculation in the schematic review catches mismatches before layout.

Oscillator tolerance matters more than most teams expect. The CAN specification requires that node oscillators be within 1.58% for reliable sampling, but many low-cost microcontrollers use internal RC oscillators with ±1-2% tolerance over temperature. I never use the internal RC for CAN on a production node; always use a quartz crystal or dedicated CAN oscillator with ±0.1% or better. I have traced winter failures on construction equipment directly to RC drift causing nodes to mis-sample near the edge of PHASE_SEG2.

// Zephyr: CAN timing configuration with sample point enforcement
// Using Zephyr CAN driver - verifies timing against controller limits
const struct device *can_dev = DEVICE_DT_GET(DT_CHOSEN(zephyr_canbus));

struct can_timing timing;
int ret = can_calc_timing(can_dev, &timing, 500000, 875);
// 500 kbps target, 87.5% sample point request
if (ret == 0) {
    printk("Computed: prescaler %u, prop_seg %u, phase_seg1 %u, phase_seg2 %u, sjw %u\n",
           timing.prescaler, timing.prop_seg, timing.phase_seg1, 
           timing.phase_seg2, timing.sjw);
    ret = can_set_timing(can_dev, &timing);
}
if (ret != 0) {
    printk("Timing rejected by controller, falling back to 80 percent sample point\n");
}

After setting timing, validate it empirically. Place two nodes at opposite ends of the longest harness, flood the bus at 80% load with a CAN stress tool, and thermal cycle the enclosure from your minimum to maximum rated temperature while monitoring error counters (TEC/REC) and retransmission counts. If you see TEC incrementing even slowly, your timing margin is insufficient — the bus is correcting but telling you it is stressed.

Error Handling, Bus-Off Recovery and Fault Confinement in Deployed Systems

CAN's fault confinement is one of its strongest features, yet firmware often ignores it until a field failure. Each CAN controller maintains two counters: Transmit Error Counter (TEC) and Receive Error Counter (REC). On detecting errors, counters increment; on successful transfers they decrement. Thresholds define states: error-active (TEC < 128, REC < 128), error-passive (128-255), and bus-off (TEC > 255). A bus-off node disconnects itself from the bus and must be recovered.

The default behavior on many controllers after bus-off is to stay off until software reinitializes. In safety-critical systems you cannot blindly auto-recover without understanding why you went bus-off. In my sensor gateways, I implement a graduated recovery: on entering bus-off, log the TEC/REC values, error frame history, and last transmitted identifier to non-volatile memory, wait 1 second, then attempt recovery. If bus-off recurs within 60 seconds, double the backoff and raise a diagnostic trouble code. After three rapid bus-off events, stay off and require external intervention. This prevents a faulty transceiver from flapping the entire bus.

Detecting Intermittent Faults Before They Cause Bus-Off

Most stacks only expose bus-off interrupts. I poll TEC/REC at 10 Hz as part of health monitoring and trend them. A node that sits at TEC = 20-40 continuously is telling you about marginal signal integrity — perhaps a connector with corrosion or a cable routed alongside a power cable. Catching that trend lets maintenance replace a harness during scheduled downtime rather than during an unplanned stoppage. Expose these counters over your diagnostic CAN messages or via a sideband interface so a plant SCADA system can alert.

// Linux SocketCAN: Monitor CAN error frames and bus state from userspace
#include <linux/can.h>
#include <linux/can/error.h>

int s = socket(PF_CAN, SOCK_RAW, CAN_RAW);
struct can_filter err_filter;
err_filter.can_id = CAN_ERR_FLAG | CAN_ERR_MASK;
err_filter.can_mask = CAN_ERR_FLAG;
setsockopt(s, SOL_CAN_RAW, CAN_RAW_FILTER, &err_filter, sizeof(err_filter));

struct can_frame frame;
while (read(s, &frame, sizeof(frame)) > 0) {
    if (frame.can_id & CAN_ERR_FLAG) {
        if (frame.data[1] & CAN_ERR_CRTL_TX_PASSIVE)
            log_warning("Node entered error-passive: TEC elevated");
        if (frame.can_id & CAN_ERR_BUSOFF)
            log_critical("Bus-off detected: data[2]=0x%02x TEC snapshot", frame.data[2]);
        if (frame.can_id & CAN_ERR_BUSERROR)
            log_error("Bit error at %02x, check termination/stubs", frame.data[3]);
    }
}

For nodes that bridge to other buses, fault propagation is a real risk. If your CAN node also aggregates data from analog sensors, noisy power or inadequate ADC decoupling can cause the MCU to brown out, leading to malformed CAN frames that drive bus errors. Techniques from ADC Sampling Techniques: Resolution, Noise Reduction and Oversampling — proper grounding, oversampling, and independent analog supplies — directly improve CAN stability by keeping the MCU in its valid operating range during transients.

Higher-Layer Protocols and Sensor Interfacing: From Raw CAN to CANopen and J1939

Raw CAN gives you arbitration and error detection, but not semantics. If you define your own payload packing — two bytes for temperature here, four bytes for pressure there — every integrator reinvents parsing and you lose interoperability. For sensor networks, I strongly recommend adopting a higher-layer protocol unless you have a compelling reason not to.

CANopen is the most structured option for industrial sensors. It defines object dictionaries, PDOs for real-time data (transmitted without protocol overhead, ideal for 10-100 Hz sensor streams), and SDOs for configuration. A CANopen temperature sensor exposes its reading at a standardized index (e.g., 0x6000 range for analog inputs), so a PLC can discover it regardless of vendor. J1939 is the equivalent in automotive/heavy equipment, built on 29-bit identifiers with Parameter Groups (PGNs) and Suspect Parameter Numbers (SPNs). If your device must talk to an engine ECU or telematics unit, J1939 is not optional.

Mapping Sensor Data Efficiently

Classical CAN limits you to 8 data bytes per frame. CAN FD extends that to 64 bytes and raises data-phase bit rates. This matters for sensor interfacing when you need to report multiple measurements atomically — for example, a 6-axis IMU with 3 accelerometer and 3 gyroscope values plus temperature. With classical CAN you would split across 2-3 frames and force the receiver to reassemble, introducing synchronization issues. With CAN FD you pack all six int16 values plus a timestamp in one frame. I've found that upgrading to CAN FD is justified even at the same arbitration bit rate if your sensor fusion requires coherent multi-axis samples.

When integrating legacy sensors that only speak raw CAN, create a gateway object dictionary on your main controller that mirrors those raw IDs into CANopen PDOs. That keeps the wire protocol efficient while presenting a clean abstraction to the application layer. Document your PGN or PDO mapping in the device EDS/DCF file so it can be imported into network configuration tools — undocumented CAN mappings are the number one time sink during commissioning, based on my field logs.

FeatureCAN 2.0 ClassicCAN FDLIN (for comparison)
Max payload8 bytes64 bytes8 bytes
Max arbitration rate1 Mbps1 Mbps (arbitration)19.2 kbps
Max data rate1 Mbps5-8 Mbps (data phase)19.2 kbps
Bus topologyMulti-master, linear bus with 120Ω terminationSame, stricter stub requirementsSingle-master, single-wire
Error handlingCRC15, ACK, fault confinement, bus-offCRC17/21, improved error detectionChecksum, no fault confinement
Typical sensor useGeneral industrial / automotive ECUs, low-rate sensorsMulti-axis sensors, firmware update over CAN, ADAS sensor fusionLow-cost body electronics: switches, seats, climate knobs
Hardware costLowest — mature transceivers and controllersModerately higher, requires FD-capable controllerLowest per node, but needs LIN master

If you are evaluating CAN FD, remember that mixing CAN 2.0 and CAN FD nodes on the same bus requires that all nodes tolerate FD frames, even if they do not transmit them. A classical node that does not understand the FD frame format will flag it as an error and drive error frames, collapsing throughput. Plan migration by either upgrading all nodes or segmenting the bus with a gateway. The FreeRTOS Documentation provides practical examples of gateway tasks that bridge classic and FD segments using separate queues.

Designing Robust CAN Nodes: Hardware Isolation, Filtering and Embedded Software Architecture

A sensor node that works on a dev board often fails in a steel cabinet with shared 24V. My hardware baseline for industrial CAN nodes includes: TVS diodes on CAN_H/CAN_L referenced to ground (e.g., PESD1CAN), a common-mode choke (e.g., ACT45B) for EMI immunity, an isolated transceiver for off-board runs, and a dedicated 5V regulator for the transceiver that can deliver peak dominant current without drooping the MCU supply. I also add a 47pF-100pF capacitor from each line to ground for ESD, placed after the choke to avoid detuning it.

On the firmware side, architecture determines whether your node can keep up. I structure CAN firmware around three priorities: a high-priority ISR that only moves frames between controller FIFOs and a lock-free ring buffer, a medium-priority task that handles real-time PDO transmission and sensor acquisition timing, and a low-priority task for SDO/configuration and diagnostics. Never parse or log inside the ISR — I have measured ISR latency spikes from 8us to 180us when teams added printf debugging to the CAN interrupt. At 500 kbps with 60% bus load, that spike guarantees FIFO overrun.

Deterministic Sensor Sampling and CAN Transmission

For periodic sensors (e.g., pressure sampled at 1 kHz, reported at 100 Hz), decouple acquisition from transmission. Sample on a hardware timer interrupt, filter/average in a task, and transmit on a separate timer that is phase-aligned to the bus schedule. If you sample and transmit in the same task, jitter in sensor readout adds jitter to the CAN schedule, which in turn adds jitter to the receiver's control loop. In one motion control application, moving from ad-hoc transmission to timer-aligned PDOs reduced velocity ripple by 18% without changing any control gains — the improvement came solely from deterministic timing.

Apply hardware acceptance filtering aggressively. For an STM32H7 with dual FDCAN, you can use range filters, dual ID filters, or classic mask filters. I prefer to dedicate one FIFO to high-priority emergency frames and another to normal data, so an emergency burst cannot overflow the data FIFO. That separation also simplifies deadline monitoring: if the emergency FIFO depth exceeds 2, you know the system is under duress and can shed non-critical transmissions.

Finally, plan for field diagnostics. Reserve one CAN ID per node for a heartbeat that includes TEC/REC, supply voltage, MCU temperature, and lost-frame counters. Make the heartbeat period configurable via SDO, and allow a bus analyzer to request detailed logs over SDO without interrupting PDO traffic. The time invested in that diagnostic surface pays back the first time you troubleshoot a 30-node installation remotely instead of sending a technician with an oscilloscope to probe the bus at each connector.

Frequently Asked Questions

Can I run CAN bus without termination resistors for short bench testing?

For a 30cm bench harness at low bit rates you can sometimes communicate with no termination, but I do not recommend it even for lab work. Unterminated stubs create reflections that cause intermittent CRC errors, and those errors teach you nothing about your final installation. Always install two 120-ohm resistors at the ends, even on the bench. If your adapter has built-in termination, account for it — two adapters with 120-ohm each plus two board resistors equals double termination and weak differential amplitude.

How many nodes can I realistically put on a single CAN segment for sensor data?

The protocol allows many identifiers, but the electrical load and bus arbitration limit you. With proper transceivers (fan-out of ~30-50 nodes per ISO 11898-2) and bus load kept under 60% average for margin, 20-30 sensor nodes per segment at 500 kbps is a practical ceiling for deterministic latency. Beyond that, segment the network with a gateway or bridge. Monitor bus load continuously — I start load-shedding or segmenting when sustained load exceeds 55% for more than 10 seconds.

When should I choose CAN FD over classical CAN for sensor interfaces?

Choose CAN FD when you need atomic multi-value sensor reports (e.g., 6-axis IMU, 3-phase power measurement, or vibration spectra) that do not fit in 8 bytes, when you need faster firmware updates over the bus, or when higher data rates can reduce bus load without raising arbitration speed. Keep in mind that CAN FD requires FD-capable controllers and stricter timing control, and mixing classical and FD nodes without verification will destabilize the bus.

What is the single most effective improvement for noisy industrial CAN installations?

Proper shield grounding and isolation. Ground the cable shield at one point only (typically at the master or cabinet entry) to avoid ground loops, use twisted pair with characteristic impedance close to 120 ohms, add a common-mode choke, and isolate nodes that leave the cabinet. After that, verify bit timing with sample points of 80-87.5% and SJW of 3-4 TQ, and enable continuous TEC/REC monitoring. That combination resolves the vast majority of noise-related field issues I encounter.

Related Articles

References & Standards: FreeRTOS Documentation · Zephyr Project Documentation · MQTT Specification