SPI Communication: Full-Duplex Data Transfer for High-Speed Sensors

When I first moved from slow environmental sensors to a 6-axis IMU streaming at 4 kHz, my tried-and-true I2C bus collapsed. The bus was saturated, the timing jitter was unacceptable, and the host was spending more cycles arbitrating than acquiring. Switching that design to SPI cut the transaction time by an order of magnitude and, more importantly, gave me simultaneous transmit and receive on every clock edge. That full-duplex capability is the core reason SPI remains the default for high-speed sensors — from MEMS gyroscopes and barometric pressure sensors to 24-bit ADCs and high-resolution encoders — where you cannot afford to waste a clock cycle. This article breaks down how to use SPI correctly for continuous, high-throughput sensor data, with the hardware details and firmware patterns I've relied on in production systems.

SPI Signal Anatomy: MOSI, MISO, SCLK and Chip Select Choreography

SPI is deceptively simple at the schematic level. You have four logical signals: Master Out Slave In (MOSI), Master In Slave Out (MISO), Serial Clock (SCLK), and Chip Select (CS), sometimes labeled SS or CSN. Unlike I2C, there is no addressing and no open-drain arbitration — the master generates the clock and asserts a dedicated CS line for each slave it wants to talk to. Data shifts out on MOSI and in on MISO at the same time, one bit per clock pulse, typically MSB-first.

In my experience, the most misunderstood pin is Chip Select. New designers treat it as a static enable. It is not. CS frames the entire transaction and controls the slave's internal state machine. Many sensors, like the Bosch BMI270 or ST LIS3DH, use the rising edge of CS to latch a command and reset their SPI shift register. If you leave CS asserted between transactions to “save time,” you will corrupt the next read. I always drive CS as a GPIO-controlled frame signal, asserted low at the start of a transfer and de-asserted high immediately after, with a minimum setup time before the first clock.

Push-Pull vs Open-Drain and Tri-State Behavior

All SPI lines are push-pull, which is why you can hit 20-30 MHz on a clean board without pull-ups. MISO is the exception: slaves must tri-state MISO when their CS is de-asserted so multiple slaves can share the bus. I've debugged boards where a cheap sensor module never released MISO, holding it low and jamming every other device. Always check the sensor datasheet for tDIS (MISO disable time) and verify with a scope that MISO floats when CS is high. If you are mixing 3.3V and 1.8V sensors, use a proper level translator; resistive dividers add capacitance and ruin your rise times at high clock rates.

Four-Wire, Three-Wire and Dual-SPI Variants

The classic four-wire implementation is full-duplex by definition. Some compact sensors offer a three-wire half-duplex mode (often called 3-wire SPI) where a single SDIO pin is bidirectional to save a pin. I avoid it for high-speed work — you lose the full-duplex advantage and add direction-switching overhead in the driver. Keep four-wire for any sensor sampling above 1 kHz. Dual and Quad SPI, where MOSI/MISO become bidirectional data lanes, are common in flash memory but rare in sensor interfaces, so I won't focus on them here.

Clock Polarity, Phase and Mode Selection From the Datasheet

Every SPI sensor specifies its clock mode using two bits: Clock Polarity (CPOL) and Clock Phase (CPHA). Together they define four modes (0-3) that determine when data is sampled and when the clock idles.

I've found that datasheets describe this in at least three different ways, which creates confusion. Some say “data is valid on the rising edge,” others say “CPOL=0, CPHA=1.” The safest approach is to map the timing diagram directly. Look at where SCLK idles (low for CPOL=0, high for CPOL=1) and whether the first data bit is set up before the first clock edge (CPHA=0) or on the first edge (CPHA=1). Most MEMS sensors use Mode 0 (CPOL=0, CPHA=0) or Mode 3 (CPOL=1, CPHA=1). In both, data is sampled on the rising edge if you squint — but the idle polarity matters for signal integrity.

A practical example: the popular Invensense ICM-42688-P defaults to Mode 0 for reads and Mode 3 for writes in some configurations, while Analog Devices ADXL355 is strictly Mode 0. If you configure the wrong mode, you will see stable but shifted data — typically off by one bit — which looks like a noisy sensor rather than a protocol error. I always validate mode by reading a known WHO_AM_I register (usually 0x75 or 0x0F) before streaming. If that read returns 0xFF or a bit-shifted value, your mode or bit order is wrong.

// STM32 HAL: Configuring SPI1 for Mode 0, 8-bit, 8 MHz
// Master, software NSS management for precise CS framing
hspi1.Instance = SPI1;
hspi1.Init.Mode = SPI_MODE_MASTER;
hspi1.Init.Direction = SPI_DIRECTION_2LINES; // Full-duplex
hspi1.Init.DataSize = SPI_DATASIZE_8BIT;
hspi1.Init.CLKPolarity = SPI_POLARITY_LOW;   // CPOL = 0
hspi1.Init.CLKPhase = SPI_PHASE_1EDGE;       // CPHA = 0 -> Mode 0
hspi1.Init.NSS = SPI_NSS_SOFT;
hspi1.Init.BaudRatePrescaler = SPI_BAUDRATEPRESCALER_8; // 64MHz / 8 = 8MHz
hspi1.Init.FirstBit = SPI_FIRSTBIT_MSB;
HAL_SPI_Init(&hspi1);

// Framed read of WHO_AM_I (register 0x75) - Mode 0 sensor
uint8_t tx[2] = { 0x75 | 0x80, 0x00 }; // MSB=1 indicates read on many sensors
uint8_t rx[2] = {0};
HAL_GPIO_WritePin(CS_GPIO_Port, CS_Pin, GPIO_PIN_RESET);
HAL_SPI_TransmitReceive(&hspi1, tx, rx, 2, HAL_MAX_DELAY);
HAL_GPIO_WritePin(CS_GPIO_Port, CS_Pin, GPIO_PIN_SET);
// rx[1] should now contain the expected ID, e.g., 0x47

Clock frequency is the other half of this equation. Your sensor will list a maximum SCLK, often 10 MHz for precision ADCs and up to 24 MHz for IMUs. That is the absolute maximum at ideal voltage and temperature. I derate by 20% for production margin and measure actual SCLK on the oscilloscope, including overshoot. A 16 MHz clock with 3 ns rise time on a 10 cm trace will ring badly if you don't have a 33-47 ohm series resistor near the master.

Achieving True Full-Duplex Throughput for High-Speed Sensor Sampling

Half-duplex protocols force you to choose: you are either sending a command or receiving data. SPI's full-duplex shift registers mean every byte you clock out returns a byte in parallel. For sensors, this is not a theoretical benefit — it is how you sustain high data rates without dead time.

Consider a typical 3-axis accelerometer configured for burst read. You send a single-byte read command (register address with R/W bit set), and as you continue clocking dummy bytes (0x00), the sensor streams acceleration data back on MISO. The master is simultaneously transmitting zeros while capturing sensor data on every edge. There is no turnaround delay. If you implement this with a blocking TransmitReceive loop at 8 MHz, a 6-byte burst (1 byte address + 6 bytes data) takes ~7 microseconds. That leaves headroom to sample at 10 kHz even on a 48 MHz Cortex-M0.

Where I see teams underutilize full-duplex is when they split transmit and receive into two separate HAL calls. That inserts an inter-byte gap where CS may glitch or the FIFO empties. Always use a single atomic transceive for a framed transaction. If you need to evaluate alternatives for slower peripherals, see the comparison in I2C Protocol Tutorial: Addressing, Clock Stretching and Multi-Master — I2C is excellent for configuration buses but it cannot match this concurrent transfer model.

Feature SPI I2C UART
Duplex Mode Full-duplex (simultaneous TX/RX) Half-duplex (shared SDA line) Full-duplex (separate TX/RX lines)
Clocking Synchronous, master-generated SCLK Synchronous, master-generated SCL with clock stretching Asynchronous, baud-rate derived
Max Sensor Bus Speed 1-50 MHz (typically 8-24 MHz) 0.1-1 MHz (3.4 MHz in high-speed mode) 0.115-3 Mbps (limited by baud error)
Slave Selection Dedicated CS per slave 7/10-bit address on bus Point-to-point only
Overhead Per Byte None (continuous streaming) ACK bit + address byte Start/stop bits per byte
Best Use Case High-speed sensor streaming, ADCs, IMUs Low-speed config, temperature, EEPROM Debug, GPS, inter-board links

For high-speed ADCs, the utilization is even more explicit. The ADC Sampling Techniques: Resolution, Noise Reduction and Oversampling flow often uses SPI to pull 24-bit conversion results while the next conversion is in progress. With a 1 MSPS ADC like the ADS8688, you assert CS, clock out the channel selection for the next sample on MOSI, and simultaneously read the previous conversion result on MISO. If you were to do this half-duplex, you would lose 50% of your theoretical throughput.

Calculating Real Bus Utilization

Your effective bandwidth is not just SCLK/8. You lose time to CS setup/hold (tCSS, tCSH), inter-frame gaps, and software latency. I budget it like this: Effective Bytes/s = SCLK / (8 * (1 + gap_ratio)). On a bare-metal loop with DMA, gap_ratio can be <0.05. With interrupt-driven byte-by-byte handling inside an RTOS, it can exceed 0.5. Measure it by toggling a debug GPIO at CS assert/deassert and use a logic analyzer to confirm you are actually transferring during 90%+ of your sampling window.

Daisy-Chaining vs Independent Slave Select for Multi-Sensor Topologies

When you have three identical sensors on one board, you have two wiring options. The straightforward approach is independent Slave Select: MISO, MOSI, and SCLK are shared, and each sensor gets its own CS line from the master. This keeps firmware simple and isolates faults — if one sensor holds MISO, you can deselect it and talk to the others.

Daisy-chaining is the alternative, supported by sensors like the ADXL355 and MAX31865. Here, you connect MOSI to the first sensor's SDI, its SDO to the next sensor's SDI, and so on, with a single shared CS and SCLK. Data ripples through the chain like a shift register. You clock out N words and clock back N words containing all sensor results. It saves GPIOs and routing, but every transaction now involves all devices, even if you only want one. I've found that daisy-chaining is elegant on paper but painful to debug when one device in the middle fails — the whole chain shifts incorrectly and you get no clean isolation.

In my experience, independent CS is almost always the right choice for high-reliability designs. You pay with two extra GPIOs, but you gain per-sensor retries and the ability to run different clock speeds per device if needed. If you are absolutely pin-constrained, consider using a 74HC138 3-to-8 decoder to generate up to 8 CS lines from 3 GPIOs, at the cost of slightly slower CS decoding. One tip: add a 10k pull-up on each CS so sensors stay deselected during MCU reset — otherwise they may interpret random SCLK toggles during boot as commands.

// Zephyr RTOS: Full-duplex transceive with independent CS control
// See Zephyr Project Documentation (https://docs.zephyrproject.org/) SPI API
struct spi_config spi_cfg = {
    .frequency = 8000000,
    .operation = SPI_WORD_SET(8) | SPI_TRANSFER_MSB | SPI_MODE_CPOL | SPI_MODE_CPHA,
    .slave = 0,
    .cs = {
        .gpio = GPIO_DT_SPEC_GET(DT_NODELABEL(spi1), cs_gpios),
        .delay = 2, // microseconds CS setup delay
    },
};

const struct device *spi_dev = DEVICE_DT_GET(DT_NODELABEL(spi1));
uint8_t tx_buf[7] = { 0x3B | 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00 };
uint8_t rx_buf[7] = {0};

struct spi_buf tx_spibuf = { .buf = tx_buf, .len = sizeof(tx_buf) };
struct spi_buf rx_spibuf = { .buf = rx_buf, .len = sizeof(rx_buf) };
struct spi_buf_set tx_set = { .buffers = &tx_spibuf, .count = 1 };
struct spi_buf_set rx_set = { .buffers = &rx_spibuf, .count = 1 };

// Atomic transceive - clocks 7 bytes out and in simultaneously
int ret = spi_transceive(spi_dev, &spi_cfg, &tx_set, &rx_set);
if (ret == 0) {
    int16_t ax = (rx_buf[1] << 8) | rx_buf[2];
    // rx_buf[0] is dummy (command echo), bytes 1-6 are sensor data
}

For topology decisions, remember signal loading. Each additional slave adds ~5-10 pF of input capacitance. At 16 MHz, three slaves and 5 cm traces may still be fine, but six slaves will round your clock edges enough to violate setup times. Buffer SCLK with a 74LVC125 if you exceed four loads.

DMA, Interrupts and Polling: Moving SPI Data Without Starving the CPU

How you move data between the SPI peripheral and memory determines whether your CPU can do anything else while sampling. I've implemented all three methods in production.

Polling is the simplest: you busy-wait on the TXE/RXNE flags for each byte. It is deterministic and easy to reason about, but it burns 100% CPU during the transfer. I use it only for one-off configuration writes or WHO_AM_I reads at boot. Never poll in a high-speed sampling loop.

Interrupt-driven transfer fires an ISR per byte or per FIFO threshold. On a 72 MHz Cortex-M4, the overhead of entering and exiting the ISR can be 0.5-1 µs, which dominates if you are streaming at 1 kHz with 14-byte frames. It also introduces jitter if a higher-priority interrupt preempts the SPI ISR.

DMA is the correct answer for continuous sensor acquisition. You configure two DMA channels — one for TX, one for RX — linked to the SPI peripheral. The CPU sets up the descriptor and sleeps or runs a control loop while the DMA engine shifts data. I pair this with a timer-triggered CS assertion to create a fully hardware-timed sampling pipeline with sub-microsecond jitter. According to the FreeRTOS Documentation, DMA completion is best signaled with a direct-to-task notification or binary semaphore from the DMA ISR, not by polling a flag in a task.

// STM32 DMA + SPI full-duplex circular acquisition (pseudo-code)
// Timer TIM3 triggers CS and DMA transfer at precise 8 kHz rate
void spi_dma_init(void) {
    // DMA1 Channel 3 for SPI1_TX, Channel 2 for SPI1_RX (STM32F4 example)
    hdma_tx.Instance = DMA1_Channel3;
    hdma_tx.Init.Direction = DMA_MEMORY_TO_PERIPH;
    hdma_tx.Init.PeriphInc = DMA_PINC_DISABLE;
    hdma_tx.Init.MemInc = DMA_MINC_ENABLE;
    hdma_tx.Init.PeriphDataAlignment = DMA_PDATAALIGN_BYTE;
    hdma_tx.Init.Mode = DMA_NORMAL;
    HAL_DMA_Init(&hdma_tx);
    __HAL_LINKDMA(&hspi1, hdmatx, hdma_tx);

    hdma_rx.Instance = DMA1_Channel2;
    hdma_rx.Init.Direction = DMA_PERIPH_TO_MEMORY;
    hdma_rx.Init.Mode = DMA_CIRCULAR; // Circular for continuous streaming
    HAL_DMA_Init(&hdma_rx);
    __HAL_LINKDMA(&hspi1, hdmarx, hdma_rx);
}

void TIM3_IRQHandler(void) {
    if (__HAL_TIM_GET_FLAG(&htim3, TIM_FLAG_UPDATE)) {
        __HAL_TIM_CLEAR_FLAG(&htim3, TIM_FLAG_UPDATE);
        HAL_GPIO_WritePin(CS_GPIO_Port, CS_Pin, GPIO_PIN_RESET);
        // Start 8-byte transceive: 1 cmd + 7 data bytes
        HAL_SPI_TransmitReceive_DMA(&hspi1, dma_tx_buffer, dma_rx_buffer, 8);
    }
}

void DMA1_Channel2_IRQHandler(void) {
    HAL_DMA_IRQHandler(&hdma_rx);
    HAL_GPIO_WritePin(CS_GPIO_Port, CS_Pin, GPIO_PIN_SET);
    BaseType_t xHigherPriorityTaskWoken = pdFALSE;
    vTaskNotifyGiveFromISR(sensorTaskHandle, &xHigherPriorityTaskWoken);
    portYIELD_FROM_ISR(xHigherPriorityTaskWoken);
}

One warning: if you share the SPI bus with multiple sensors, DMA descriptors must be reconfigured for each CS target. Do not attempt to have two DMA-driven sensors overlapping on the same bus. I use a mutex around the SPI bus and a dedicated bus manager task that serializes DMA requests.

Signal Integrity and Layout Pitfalls at 10 MHz and Beyond

SPI looks slow compared to USB or Ethernet, but at 16-24 MHz your traces behave like transmission lines. I've chased phantom bit errors for days only to find ringing on SCLK that double-clocked a sensor.

My layout rules for reliable high-speed SPI are now non-negotiable: Keep SCLK, MOSI, and MISO length-matched within 5 mm and under 50 mm total length if possible. Route them as a group over a continuous ground plane — never cross a plane split. Place a 33-ohm series resistor at the source (MCU side) for SCLK and MOSI to damp overshoot; most logic analyzers won't show this, but a 200 MHz scope will. Avoid stubs: if you have a multi-drop bus with independent CS, use a star topology from the master, not a chain that snakes through each sensor connector.

Ground bounce is another silent killer. When you have a high-current load (like a motor) sharing ground return with your sensor SPI ground, the shift in ground reference can violate VIL/VIH at the sensor input. I isolate analog sensor ground from power ground with a single-point connection near the ADC or use a dedicated LDO for the sensor rail.

Finally, don't ignore the sensor's own requirements. Many high-speed sensors require decoupling capacitors placed <2 mm from their VDD pin, otherwise supply droop during conversion imprints directly as noise on the SPI output drivers. Combined with proper ADC Sampling Techniques: Resolution, Noise Reduction and Oversampling at the firmware level, clean power and layout will make the difference between 12-bit and 14-bit effective resolution on the same silicon.

For debugging, a logic analyzer that decodes SPI at your target clock is essential. Tools like Saleae or sigrok can capture MOSI and MISO simultaneously and show you the actual full-duplex exchange. When timing is suspect, I also use UART Serial Communication: Baud Rate, Flow Control and Debugging as a side-channel to print filtered sensor values without disturbing the SPI bus. UART may be slower, but its asynchronous nature makes it a non-intrusive diagnostic path.

Frequently Asked Questions

Can I use SPI and I2C on the same pins with different sensors?

Not simultaneously on the same physical pins. While some MCUs allow pin multiplexing, SPI and I2C have conflicting electrical requirements — SPI is push-pull and I2C is open-drain with pull-ups. You can share the bus time-division style if you use separate peripherals and tristate unused drivers, but in practice I dedicate separate pin groups. Keeping SPI for high-speed streaming and I2C for slow configuration sensors avoids contention and simplifies drivers.

Why does my SPI sensor return 0xFF or 0x00 for every register?

All 0xFF or 0x00 usually means MISO is not driven. Check four things in order: 1) Is CS actually asserted low during the transfer (probe it), 2) Is the clock mode matched to the datasheet (try Mode 0 vs Mode 3), 3) Is the read bit set correctly in the address byte — many sensors require bit 7 = 1 for read, 4) Is the sensor powered and out of reset (check that its internal LDO has started). In my experience, a missing bitwise OR with 0x80 on the register address is the most common firmware bug.

How fast can I realistically clock SPI for continuous sensor data?

On a Cortex-M4/M7 with DMA, 10-16 MHz is reliably sustainable for continuous acquisition if your layout is clean and your sensor supports it. Above 20 MHz, you need very short traces, series termination, and often need to reduce clock for reads vs writes — many sensors allow faster writes than reads. I cap production designs at 80% of the sensor's rated maximum SCLK to account for voltage and temperature variation, and I always verify with an oscilloscope at the sensor pin, not just at the MCU pin.

Do I need to use DMA for high-speed SPI sensors?

For sampling rates above 2 kHz or frame sizes larger than 8 bytes, yes. Without DMA, interrupt latency and task preemption cause jitter and occasional overruns. DMA lets the hardware handle the bit-shifting while the CPU processes the previous sample. For lower rates (e.g., 100 Hz temperature sensor), interrupt or even polling is acceptable and simpler to maintain.

Related Articles

References & Standards: FreeRTOS Documentation · Zephyr Project Documentation · MQTT Specification