STM32 HAL Drivers: GPIO, UART, SPI and DMA Configuration

Working with STM32 HAL drivers feels straightforward until you move beyond blinking an LED. The abstraction promises portability across the entire STM32 family, but in practice, getting GPIO, UART, SPI, and DMA to cooperate reliably requires an understanding of what the HAL is doing underneath and where it tends to hide important details. Over several years of building sensor nodes, motor controllers, and industrial interfaces on F4, G4, and H7 devices, I've learned to treat CubeMX as a starting point, not a final answer. This article walks through how I configure and harden these four core peripherals for production firmware, with working code patterns and the trade-offs I actually weigh on real hardware.

GPIO Under HAL: Output Speed, Pull Resistors, and EXTI Gotchas

In my experience, most GPIO problems come from leaving CubeMX defaults untouched. The HAL structures expose a lot of useful control, but the generated code often picks a low output speed and no pull resistor, which is fine for an LED but not for a fast SPI chip select or a button that floats when unpressed. On STM32, the GPIO port is configured through GPIO_InitTypeDef, and the two fields I audit first are Speed and Pull.

Speed does not set a precise slew rate; it selects a drive strength configuration. For a push-pull output driving a short trace to an onboard sensor, GPIO_SPEED_FREQ_LOW keeps EMI down and is perfectly adequate. For an off-board SPI bus at 10 MHz or a fast switching pin that toggles in an interrupt, I step up to GPIO_SPEED_FREQ_VERY_HIGH. If you scope a LOW speed pin at 8 MHz SPI SCK, you will see rounded edges that violate the slave's setup time. I learned that on a custom board where the SD card would intermittently fail CRC - switching the SCK and MOSI pins to very high speed fixed it without any other changes.

Input Configuration and Debouncing Reality

For inputs, never rely on a floating state. If your schematic has an external pull-up, set GPIO_PULLUP to GPIO_NOPULL in HAL to avoid a voltage divider. If not, use the internal pull. For buttons and interrupt lines, the STM32 EXTI system is powerful but shares interrupt vectors. On an F401, for example, EXTI lines 5 to 9 share EXTI9_5_IRQn and lines 10 to 15 share EXTI15_10_IRQn. That means your HAL callback must check which pin triggered.

I also avoid doing any real work inside HAL_GPIO_EXTI_Callback. In my projects, that callback just sets a volatile flag or gives a FreeRTOS semaphore. Debouncing belongs in a task, not in the ISR. According to the FreeRTOS Documentation, giving a semaphore from an ISR with xSemaphoreGiveFromISR() is the safe way to wake a task that can then sample the pin 20 ms later.

// GPIO EXTI example: PA0 as interrupt, PC13 as output
void MX_GPIO_Init(void)
{
  GPIO_InitTypeDef GPIO_InitStruct = {0};

  __HAL_RCC_GPIOA_CLK_ENABLE();
  __HAL_RCC_GPIOC_CLK_ENABLE();

  // PC13 - Onboard LED as push-pull output
  GPIO_InitStruct.Pin = GPIO_PIN_13;
  GPIO_InitStruct.Mode = GPIO_MODE_OUTPUT_PP;
  GPIO_InitStruct.Pull = GPIO_NOPULL;
  GPIO_InitStruct.Speed = GPIO_SPEED_FREQ_LOW;
  HAL_GPIO_Init(GPIOC, &GPIO_InitStruct);

  // PA0 - User button with interrupt on falling edge
  GPIO_InitStruct.Pin = GPIO_PIN_0;
  GPIO_InitStruct.Mode = GPIO_MODE_IT_FALLING;
  GPIO_InitStruct.Pull = GPIO_PULLUP;
  HAL_GPIO_Init(GPIOA, &GPIO_InitStruct);

  HAL_NVIC_SetPriority(EXTI0_IRQn, 5, 0);
  HAL_NVIC_EnableIRQ(EXTI0_IRQn);
}

void HAL_GPIO_EXTI_Callback(uint16_t GPIO_Pin)
{
  if (GPIO_Pin == GPIO_PIN_0) {
    BaseType_t xHigherPriorityTaskWoken = pdFALSE;
    xSemaphoreGiveFromISR(xButtonSemaphore, &xHigherPriorityTaskWoken);
    portYIELD_FROM_ISR(xHigherPriorityTaskWoken);
  }
}

One last habit: after configuring GPIO, I lock critical pins like safety interlocks with HAL_GPIO_LockPin() during initialization. Once locked, the configuration cannot be changed until a reset, which prevents a runaway pointer from accidentally reconfiguring an enable line. It is a small safeguard that has saved a prototype during development.

UART Configuration Beyond Baud Rate: Clock Sources, FIFOs, and Error Flags

CubeMX makes UART look like just baud rate, word length, and stop bits, but clock selection is where many intermittent failures originate. On G4 and H7 parts, the UART can be clocked from PCLK, SYSCLK, LSE, or HSI16. If you run the core clock through dynamic scaling or stop modes, a UART tied to PCLK will change baud rate unexpectedly. For any UART that must stay alive during low-power modes or with precise baud timing, I explicitly select HSI16 or LSE as the kernel clock in CubeMX under Clock Configuration and verify the actual baud error stays under 2% in the generated usart.c.

Handling Overrun, Framing, and Noise Errors

By default, HAL's polling and interrupt drivers disable error interrupts after init. If you enable them correctly, you can catch overrun errors that otherwise silently corrupt data. I've found that enabling UART_IT_ERR and checking HAL_UART_ERROR_ORE in HAL_UART_ErrorCallback reveals when your DMA or interrupt handling is too slow. On a 921600 baud GPS stream on an F4, we were losing bytes until we enabled FIFO mode on the H7 variant and set thresholds properly.

For parts with FIFOs (H7, G4), configure FIFO thresholds. A UART with 16-byte TX/RX FIFOs can trigger an interrupt when the RX FIFO is half full, reducing interrupt frequency by 8x. The HAL API for this is HAL_UARTEx_SetRxFifoThreshold() and HAL_UARTEx_EnableFifoMode(). If you ignore FIFO settings, the HAL defaults to threshold 1/8, which is fine but not optimal for high throughput.

Choosing Between Polling, Interrupt, and DMA

For debug console prints, polling with HAL_UART_Transmit() and a 100 ms timeout is acceptable. For anything that runs while the main loop does other work, I use interrupt or DMA. Interrupt mode with a circular buffer works well for command-line interfaces where messages are short and irregular. For continuous streams like Modbus or telemetry, DMA is far more efficient. When you need a robust industrial link that bridges UART sensors to Ethernet, the architecture in Raspberry Pi as Industrial IoT Gateway: Modbus, MQTT and Edge Processing shows how those UART streams are typically aggregated and forwarded, which informs how I frame packets on the STM32 side to be MQTT-ready from the start, following the MQTT Specification for payload structure.

// UART2 with DMA for continuous RX using circular mode and idle line detection
UART_HandleTypeDef huart2;
DMA_HandleTypeDef hdma_usart2_rx;
uint8_t uart_rx_buffer[256];

void MX_USART2_UART_Init(void)
{
  huart2.Instance = USART2;
  huart2.Init.BaudRate = 115200;
  huart2.Init.WordLength = UART_WORDLENGTH_8B;
  huart2.Init.StopBits = UART_STOPBITS_1;
  huart2.Init.Parity = UART_PARITY_NONE;
  huart2.Init.Mode = UART_MODE_TX_RX;
  huart2.Init.HwFlowCtl = UART_HWCONTROL_NONE;
  huart2.Init.OverSampling = UART_OVERSAMPLING_16;
  huart2.Init.OneBitSampling = UART_ONE_BIT_SAMPLE_DISABLE;
  huart2.Init.ClockPrescaler = UART_PRESCALER_DIV1;
  huart2.AdvancedInit.AdvFeatureInit = UART_ADVFEATURE_NO_INIT;
  HAL_UART_Init(&huart2);

  // Enable FIFO mode on supported devices
  HAL_UARTEx_SetTxFifoThreshold(&huart2, UART_TXFIFO_THRESHOLD_1_8);
  HAL_UARTEx_SetRxFifoThreshold(&huart2, UART_RXFIFO_THRESHOLD_1_2);
  HAL_UARTEx_EnableFifoMode(&huart2);
}

void HAL_UART_MspInit(UART_HandleTypeDef* huart)
{
  if (huart->Instance == USART2) {
    __HAL_RCC_USART2_CLK_ENABLE();
    __HAL_RCC_DMA1_CLK_ENABLE();
    
    hdma_usart2_rx.Instance = DMA1_Stream5;
    hdma_usart2_rx.Init.Request = DMA_REQUEST_USART2_RX;
    hdma_usart2_rx.Init.Direction = DMA_PERIPH_TO_MEMORY;
    hdma_usart2_rx.Init.PeriphInc = DMA_PINC_DISABLE;
    hdma_usart2_rx.Init.MemInc = DMA_MINC_ENABLE;
    hdma_usart2_rx.Init.PeriphDataAlignment = DMA_PDATAALIGN_BYTE;
    hdma_usart2_rx.Init.MemDataAlignment = DMA_MDATAALIGN_BYTE;
    hdma_usart2_rx.Init.Mode = DMA_CIRCULAR;
    hdma_usart2_rx.Init.Priority = DMA_PRIORITY_LOW;
    HAL_DMA_Init(&hdma_usart2_rx);
    __HAL_LINKDMA(huart, hdmarx, hdma_usart2_rx);

    // Enable idle line interrupt to detect end of variable-length frame
    __HAL_UART_ENABLE_IT(huart, UART_IT_IDLE);
    HAL_NVIC_SetPriority(USART2_IRQn, 5, 0);
    HAL_NVIC_EnableIRQ(USART2_IRQn);
    HAL_UART_Receive_DMA(huart, uart_rx_buffer, sizeof(uart_rx_buffer));
  }
}

void USART2_IRQHandler(void)
{
  if (__HAL_UART_GET_FLAG(&huart2, UART_FLAG_IDLE)) {
    __HAL_UART_CLEAR_IDLEFLAG(&huart2);
    // Compute received length: BUFFER_SIZE - NDTR
    uint16_t remaining = __HAL_DMA_GET_COUNTER(&hdma_usart2_rx);
    uint16_t received = sizeof(uart_rx_buffer) - remaining;
    // Process frame in task, not here
  }
  HAL_UART_IRQHandler(&huart2);
}

SPI Master Configuration: Clock Polarity, Prescalers, and Manual Chip Select

SPI looks simple - four wires - but HAL exposes decisions that directly affect compatibility with sensors and flash chips. The two most common failures I see are mismatched CPOL/CPHA and incorrect NSS handling. Many devices specify mode 0 (CPOL=0, CPHA=0) while SD cards and some displays require mode 0 or 3 depending on transaction. Always check the slave datasheet timing diagram; do not assume mode 0.

For the prescaler, I calculate the SCK frequency from the APB clock, not the core clock. On an F407 running at 168 MHz, APB2 is 84 MHz. A SPI_BAUDRATEPRESCALER_8 gives 10.5 MHz SCK, which is near the limit for many breadboard wires. I start at a conservative prescaler like 32 or 64 during bring-up, verify with a scope, then increase speed. I've found that signal integrity, not the HAL setting, is often the limit.

Why I Use Software NSS Management

The HAL offers SPI_NSS_HARD_OUTPUT and SPI_NSS_SOFT. Hardware NSS can be useful for single-slave setups, but for multi-slave buses or when you need to control timing around the transaction (like holding CS low across multiple HAL calls), software NSS is more predictable. I set NSS = SPI_NSS_SOFT, configure the CS pin as a normal GPIO output, and drive it manually. This avoids a known quirk where hardware NSS on some F1/F4 revisions pulses between bytes when using DMA.

Another detail is data size. Most sensors use 8-bit frames, but some 16-bit ADCs and displays expect 16-bit data size. The HAL field DataSize changes how the DR register is accessed. If you set 16-bit mode but send an 8-bit buffer via HAL_SPI_Transmit, you will get shifted data. I keep the SPI handle at 8-bit and handle 16-bit sensors by sending two bytes manually to keep DMA alignment simple.

SPI Transaction Structure

Rather than calling HAL_SPI_TransmitReceive inline everywhere, I wrap a transaction function that asserts CS, performs the transfer, and deasserts CS with proper delays. Some slaves require a few microseconds between CS assertion and the first clock - a short __NOP() loop or a timer delay satisfies that without adding complexity. If you are evaluating whether STM32 is even the right platform for your peripheral mix, Microcontroller Selection: ARM Cortex-M, RISC-V and AVR Decision Criteria provides a structured way to compare pin counts, DMA channels, and SPI instances before you commit to a part.

Linking DMA to UART and SPI for Zero-Copy Transfers

DMA is where HAL either shines or frustrates you. The concept is clear: a DMA stream moves bytes between peripheral and memory without CPU involvement. The implementation details - stream assignment, request mapping, and interrupt priority - are where boards fail to boot or hang under load. In my experience, you should assign DMA streams before laying out the PCB for peripherals that will use them heavily, because some STM32 families have fixed DMA request mappings. On G0, you have DMA channels tied to specific peripherals; on F4/H7, you have DMA1/DMA2 streams with a multiplexer (DMAMUX).

Two configuration choices matter most: mode and increment. For UART RX circular logging, use DMA_CIRCULAR so the buffer wraps automatically. For SPI TX of a display framebuffer, use DMA_NORMAL so the transfer stops at the end and you get a transfer-complete interrupt. Peripheral increment must always be disabled (DMA_PINC_DISABLE), memory increment enabled (DMA_MINC_ENABLE) for byte arrays. Misconfiguring alignment causes hard faults - if you set DMA_PDATAALIGN_HALFWORD for an 8-bit UART, the DMA will read two bytes at a time and corrupt memory.

Double Buffering and Cache Coherency on H7

On Cortex-M7 parts like the H7, cache coherency is a real problem. The DMA writes directly to SRAM, bypassing the D-Cache. If your buffer lives in cached DTCM or AXI SRAM, the CPU may read stale cache lines. I place DMA buffers in non-cached SRAM (like SRAM2 on H7) or align them to 32 bytes and manually invalidate the cache after transfer: SCB_InvalidateDCache_by_Addr((uint32_t*)buffer, size). I have chased ghost bytes for hours before remembering this step. For M4 devices without cache, you can ignore this.

For high-speed SPI to a display or external flash, I use DMA with HAL_SPI_Transmit_DMA and rely on the HAL_SPI_TxCpltCallback to release the CS line and signal the task. Never poll on DMA completion in a tight loop; let the callback do the work. This keeps the CPU free to handle other sensors, similar to how the ESP32 for IoT: WiFi, Bluetooth and Deep Sleep for Battery-Powered Projects offloads work to keep power low, but on STM32 we do it for deterministic timing rather than battery life.

// SPI1 master with DMA for TX - manual CS control
SPI_HandleTypeDef hspi1;
DMA_HandleTypeDef hdma_spi1_tx;
uint8_t spi_tx_buffer[512];

void MX_SPI1_Init(void)
{
  hspi1.Instance = SPI1;
  hspi1.Init.Mode = SPI_MODE_MASTER;
  hspi1.Init.Direction = SPI_DIRECTION_2LINES;
  hspi1.Init.DataSize = SPI_DATASIZE_8BIT;
  hspi1.Init.CLKPolarity = SPI_POLARITY_LOW;
  hspi1.Init.CLKPhase = SPI_PHASE_1EDGE;
  hspi1.Init.NSS = SPI_NSS_SOFT;
  hspi1.Init.BaudRatePrescaler = SPI_BAUDRATEPRESCALER_16;
  hspi1.Init.FirstBit = SPI_FIRSTBIT_MSB;
  hspi1.Init.TIMode = SPI_TIMODE_DISABLE;
  hspi1.Init.CRCCalculation = SPI_CRCCALCULATION_DISABLE;
  HAL_SPI_Init(&hspi1);
}

void HAL_SPI_MspInit(SPI_HandleTypeDef* hspi)
{
  if (hspi->Instance == SPI1) {
    __HAL_RCC_SPI1_CLK_ENABLE();
    __HAL_RCC_DMA1_CLK_ENABLE();

    hdma_spi1_tx.Instance = DMA1_Stream3;
    hdma_spi1_tx.Init.Request = DMA_REQUEST_SPI1_TX;
    hdma_spi1_tx.Init.Direction = DMA_MEMORY_TO_PERIPH;
    hdma_spi1_tx.Init.PeriphInc = DMA_PINC_DISABLE;
    hdma_spi1_tx.Init.MemInc = DMA_MINC_ENABLE;
    hdma_spi1_tx.Init.PeriphDataAlignment = DMA_PDATAALIGN_BYTE;
    hdma_spi1_tx.Init.MemDataAlignment = DMA_MDATAALIGN_BYTE;
    hdma_spi1_tx.Init.Mode = DMA_NORMAL;
    hdma_spi1_tx.Init.Priority = DMA_PRIORITY_MEDIUM;
    HAL_DMA_Init(&hdma_spi1_tx);
    __HAL_LINKDMA(hspi, hdmatx, hdma_spi1_tx);

    HAL_NVIC_SetPriority(DMA1_Stream3_IRQn, 6, 0);
    HAL_NVIC_EnableIRQ(DMA1_Stream3_IRQn);
  }
}

void SPI_TransmitBuffer_DMA(uint8_t *data, uint16_t len)
{
  HAL_GPIO_WritePin(GPIOA, GPIO_PIN_4, GPIO_PIN_RESET); // CS low
  HAL_SPI_Transmit_DMA(&hspi1, data, len);
}

void HAL_SPI_TxCpltCallback(SPI_HandleTypeDef *hspi)
{
  if (hspi->Instance == SPI1) {
    HAL_GPIO_WritePin(GPIOA, GPIO_PIN_4, GPIO_PIN_SET); // CS high
    BaseType_t xWoken = pdFALSE;
    xSemaphoreGiveFromISR(xSpiDoneSemaphore, &xWoken);
    portYIELD_FROM_ISR(xWoken);
  }
}

Interrupt Priorities, HAL Callbacks, and Coexistence with an RTOS

One of the most underrated configuration steps is NVIC priority grouping. HAL uses SysTick for HAL_Delay and HAL_GetTick, and it also relies on interrupt priorities for DMA and peripheral IRQs. If you run FreeRTOS, you must configure priorities to respect configMAX_SYSCALL_INTERRUPT_PRIORITY. I set the priority grouping to 4 bits of preemption (NVIC_PRIORITYGROUP_4) and ensure no peripheral ISR that calls a FreeRTOS API runs at a priority numerically lower (higher urgency) than configMAX_SYSCALL_INTERRUPT_PRIORITY. In practice that means DMA and UART interrupts sit at 5-6, while critical faults remain at 0-1.

Callback Discipline

HAL's callback model - HAL_UART_TxCpltCallback, HAL_SPI_RxCpltCallback, HAL_UART_ErrorCallback - is convenient but easy to misuse. These callbacks run in interrupt context. I keep them short: clear flags, give a semaphore, or set a flag. Heavy processing like parsing a Modbus frame or writing to flash should happen in a task. I also avoid calling blocking HAL functions from callbacks. Calling HAL_UART_Transmit with a timeout inside HAL_UART_RxCpltCallback will deadlock if that UART's interrupt is the same one you are already in.

For timing, I replace HAL_Delay with RTOS delays once the scheduler starts. HAL_Delay depends on SysTick incrementing via HAL_IncTick, which FreeRTOS may also use. If you configure FreeRTOS to use a different tick source, call HAL_IncTick from the SysTick handler manually, or the HAL timeout logic will break. The Zephyr Project Documentation discusses similar tick management issues when abstracting HALs across RTOSes, and the same principles apply even if you stay with FreeRTOS.

Validating HAL Configurations with a Logic Analyzer and Hard Fault Analysis

I've learned not to trust that CubeMX output works until I see it on a scope or logic analyzer. A $10 Salae clone with PulseView will confirm SCK frequency, CPOL/CPHA idle levels, and CS timing. For UART, I verify that the idle line time matches the expected frame gap, especially when using DMA idle detection. One project had a UART that worked at room temperature but failed at -20C because the HSI clock drifted and the baud error exceeded 3.5% - switching the UART kernel clock to HSE fixed it, and the analyzer showed the stretched bit time clearly.

Common HAL Pitfalls I Check First

When a peripheral does not start, I check three things in order: clocks, GPIO alternate function, and DMA stream assignment. HAL's MspInit functions are where clocks and GPIO AF are enabled. If you rename a pin in CubeMX but forget to update the AF mapping (e.g., PA9/PA10 for USART1 is AF7, but PB6/PB7 is also AF7 on some parts), the pin will stay in GPIO mode and the UART will transmit into nothing. The DMA request number is equally silent - an incorrect DMA_REQUEST_* value compiles but never triggers a transfer, and HAL_DMA_Init returns OK.

Hard faults from DMA usually point to misaligned buffers or accessing a peripheral while its clock is disabled. Enable the Memory Protection Unit or at least hard fault handlers that dump HFSR, CFSR, and BFAR over SWO. I keep a minimal fault handler that prints these registers via ITM - it immediately tells me if the fault was a precise bus error from DMA writing to an invalid address. This level of verification is what makes HAL reliable for production instead of just prototyping, and it bridges well to decisions you will make when scaling to more complex platforms discussed in RISC-V Microcontrollers: An Open Architecture for Embedded Systems, where the driver model differs but the validation workflow is identical.

Transfer Method CPU Load Latency Best Use Case HAL API Example
Polling High - blocks task Deterministic, but stalls Simple debug prints, bring-up HAL_UART_Transmit() with timeout
Interrupt Medium - per byte ISR Low for short messages CLI, irregular command frames HAL_UART_Receive_IT() + Callback
DMA Normal Very low - setup only Medium - setup overhead SPI display, flash write, bulk TX HAL_SPI_Transmit_DMA()
DMA Circular + Idle Very low Low - idle interrupt Continuous UART RX, sensor streams HAL_UART_Receive_DMA() + IDLE

Frequently Asked Questions

When should I use DMA with HAL UART instead of interrupts?

Use DMA when you expect continuous or high-throughput data, typically above 115200 baud with sustained traffic, or when the CPU has other real-time tasks. Interrupts work well for short, infrequent messages like AT commands. In my experience, if your UART RX interrupt fires more than 5-10% of CPU time at your baud rate, moving to DMA circular mode with idle line detection will reduce load and prevent overrun errors. DMA also pairs better with low-power designs where you want the CPU sleeping while bytes arrive.

Why does HAL_SPI hang on HAL_SPI_Transmit with timeout?

The most common cause is the SPI peripheral waiting for a clock that never starts because NSS is misconfigured or the slave is holding MISO low and the SPI is in the wrong mode. Check that you are using software NSS correctly and that CPOL/CPHA matches the slave. Another cause is calling the blocking API from inside an interrupt that has the same or higher priority than the SPI interrupt - the HAL timeout relies on SysTick, which may not increment if interrupts are masked. Use DMA or interrupt mode for SPI inside RTOS tasks.

How do I handle cache coherency for DMA buffers on STM32H7?

On Cortex-M7 devices, place DMA buffers in non-cached memory or manage cache manually. I declare buffers with __attribute__((section(".sram2"))) where SRAM2 is configured as non-cacheable via MPU, or I align buffers to 32 bytes and call SCB_InvalidateDCache_by_Addr() after DMA RX completes and SCB_CleanDCache_by_Addr() before DMA TX starts. Failure to do this results in the CPU reading stale cached data or the DMA sending stale memory contents. F4 and G4 devices do not have this issue.

Can I mix HAL with direct register access for time-critical GPIO?

Yes, and I do it frequently for bit-banging or cycle-accurate toggling. HAL is fine for initialization, but for toggling a pin at MHz rates inside a tight loop, use GPIOA->BSRR = GPIO_PIN_5 to set and GPIOA->BSRR = GPIO_PIN_5 << 16 to reset. This is a single-cycle write versus HAL_GPIO_WritePin which adds branching. Keep HAL for setup and all other code for portability, and isolate the direct register writes to a small inline function so the rest of the codebase remains readable.

Related Articles

References & Standards: FreeRTOS Documentation · Zephyr Project Documentation · MQTT Specification