Skip to content

{bp-19506} arch/arm/stm32h7: invalidate dcache after aligned SDMMC RX DMA - #19640

Merged
xiaoxiang781216 merged 1 commit into
apache:releases/13.0from
jerpelea:bp-19506
Aug 3, 2026
Merged

{bp-19506} arch/arm/stm32h7: invalidate dcache after aligned SDMMC RX DMA#19640
xiaoxiang781216 merged 1 commit into
apache:releases/13.0from
jerpelea:bp-19506

Conversation

@jerpelea

@jerpelea jerpelea commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

Summary

The aligned direct-DMA receive path only invalidated the destination buffer before the transfer in stm32_dmarecvsetup(). On the Cortex-M7 the cache can speculatively prefetch into that cacheable buffer between the pre-DMA invalidate and DMA completion, leaving stale lines that shadow the data just written by the IDMA, so the CPU reads a previously cached sector instead of the freshly received data.

Invalidate again in stm32_recvdma() once the aligned transfer completes, before the buffer is consumed. The buffer and length are cache-line aligned on this path, so no adjacent memory is affected.

This matches the STM32 AN4839 guidance that a cache invalidate is required after DMA completion and before the CPU reads the updated region, not only before the transfer starts. A related instance of the same "invalidate too early" defect on STM32H7 SPI DMA is tracked in #11594.

Root-caused on a PX4 FMUv6C (STM32H743) board where MAVLink ULog downloads were intermittently corrupted: forensic diffing showed corrupted windows were exactly 32 bytes (the D-cache line size), cache-line aligned, and byte-for-byte equal to the previous 512-byte SD sector cached in the FAT single-sector buffer. Disabling the D-cache made the corruption disappear, isolating the defect to cache coherency. After this fix, downloaded files matched the source file byte-for-byte (sha256 identical) across a 5.8 MB log spanning thousands of sectors.

Impact

RELEASE

Testing

CI

The aligned direct-DMA receive path only invalidated the destination
buffer before the transfer in stm32_dmarecvsetup(). On the Cortex-M7
the cache can speculatively prefetch into that cacheable buffer
between the pre-DMA invalidate and DMA completion, leaving stale
lines that shadow the data just written by the IDMA, so the CPU
reads a previously cached sector instead of the freshly received
data.

Invalidate again in stm32_recvdma() once the aligned transfer
completes, before the buffer is consumed. The buffer and length are
cache-line aligned on this path, so no adjacent memory is affected.

This matches the STM32 AN4839 guidance that a cache invalidate is
required after DMA completion and before the CPU reads the updated
region, not only before the transfer starts. A related instance of
the same "invalidate too early" defect on STM32H7 SPI DMA is tracked
in apache#11594.

Root-caused on a PX4 FMUv6C (STM32H743) board where MAVLink ULog
downloads were intermittently corrupted: forensic diffing showed
corrupted windows were exactly 32 bytes (the D-cache line size),
cache-line aligned, and byte-for-byte equal to the previous 512-byte
SD sector cached in the FAT single-sector buffer. Disabling the
D-cache made the corruption disappear, isolating the defect to cache
coherency. After this fix, downloaded files matched the source file
byte-for-byte (sha256 identical) across a 5.8 MB log spanning
thousands of sectors.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Yang-Rui Li <yang77567789@gmail.com>
@github-actions github-actions Bot added Arch: arm Issues related to ARM (32-bit) architecture Size: S The size of the change in this PR is small labels Aug 3, 2026
@xiaoxiang781216
xiaoxiang781216 merged commit 573c32f into apache:releases/13.0 Aug 3, 2026
16 of 26 checks passed
@jerpelea
jerpelea deleted the bp-19506 branch August 3, 2026 14:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Arch: arm Issues related to ARM (32-bit) architecture Size: S The size of the change in this PR is small

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants