Skip to content

SIMD and vectorization: auto-vectorization conditions, intrinsics, and CPU dispatch ​

One instruction, multiple data ​

SIMD (Single Instruction, Multiple Data) is the main lever contemporary CPUs use to raise compute throughput: one instruction processes multiple data elements at the same time. The wider the register, the more data per instruction:

  • SSE: 128-bit registers, 4 floats (or 2 doubles) at a time.
  • AVX/AVX2: 256-bit, 8 floats at a time.
  • AVX-512: 512-bit, 16 floats at a time.

In theory, converting a scalar loop to AVX2 multiplies throughput by ×8 outright (each instruction does arithmetic on 8 floats). We measured a dot product to see the real number:

text
===== SIMD 点积(4M float = 16MB,30 次平均)=====
  标量 dot_scalar                958.2 us
  手写 AVX2 dot_avx2              48.1 us
  AVX2/标量 = 19.9x

Nearly 20×. That number is above the theoretical 8× because the hand-written version stacks three things: SIMD width, FMA (one instruction doing the work of a multiply plus an add), and multiple accumulators (breaking the dependency chain, as covered in 04-02). First, let's fix a piece of intuition: the scalar dot_scalar at -O3 -mavx2 is in fact already auto-vectorized (g++ -O3 -mavx2 -fopt-info-vec-optimized reports loop vectorized using 32 byte vectors) — it is not "not vectorized". Then why is it still 20× slower? Because its accumulation forms a dependency chain (the assembly below shows it is a scalar horizontal accumulation chain), while the hand-written dot_avx2 has 4 accumulators running in parallel. That 20× is mostly the ILP gap of "multiple accumulators vs a single accumulator", not "vectorized vs not". That is exactly the point this article exists to make.

Auto-vectorization: the conditions + a common misconception about FP reductions ​

Under -O2 / -O3 -ftree-vectorize, the compiler tries to turn loops into SIMD automatically. The preconditions:

  1. An estimable iteration count (the compiler needs a rough idea of how many iterations there will be before it can lay out SIMD).
  2. No pointer aliasing (the compiler must be able to assume the memory accessed by two loop iterations does not overlap; __restrict helps it, see ch04-04).
  3. No function calls (or calls that inlining can eliminate).
  4. Simple branches (no complex control flow).

A fifth condition — FP reductions — deserves an update to the old wisdom. The heart of a dot product is acc += a[i] * b[i], which is an FP reduction. Vectorizing it means splitting the accumulation into multiple parallel lanes, which changes the association order of the floating-point additions, and floating-point addition is not associative ((a+b)+c ≠ a+(b+c), the mantissa rounding errors differ). Older textbooks often say "default -O3 strictly respects FP semantics, dares not change the association order, and therefore will not vectorize FP reductions" — that claim is outdated on modern GCC: since GCC 10, GCC vectorizes FP reductions in ordered (order-preserving) reduction mode, which keeps the observable accumulation order (the result is bit-for-bit identical) and uses SIMD internally only for order-preserving partial sums. Verify it yourself: compile this article's dot_scalar with g++ -O3 -mavx2 -fopt-info-vec-optimized and you will see loop vectorized using 32 byte vectors.

Then why is the scalar dot product still nearly 20× slower than the hand-written dot_avx2? Look at the assembly ordered mode actually generates: under -O3 -mavx2, the multiplies in dot_scalar are vectorized with vmulps (8-way SIMD), but the order-preserving accumulation is a string of scalar vaddss (adding horizontally, lane by lane, into the same xmm0), forming a scalar addition dependency chain. That is the price of ordered mode: it vectorizes the multiplies but cannot parallelize the accumulation (parallel accumulation would change the FP association order). Only the 4 vector accumulators of the hand-written dot_avx2 break through that limit and let 4 FMAs truly run in parallel. So that 20× is mostly the gap of "4 accumulators in parallel vs sequential scalar accumulation", not "vectorized vs not vectorized". To get the compiler to split into multiple lanes itself, you must relax FP semantics with -ffast-math (or the narrower -fassociative-math). With -ffast-math on, GCC really does rearrange dot_scalar into 4 vector accumulators (the assembly shows vfmadd231ps × 4 + unroll 32, the same shape as the hand-written dot_avx2); but measured on this machine's WSL2 the runtime barely changes, which says the 20× gap on this simple loop comes not only from the accumulator count — dot_scalar is also inlined into the REP outer loop and sits in a different measurement structure than dot_avx2, which is called as a standalone function. Integer reductions have no associativity problem, get multiple lanes by default, and are auto-vectorized more aggressively.

To find out whether your loop got vectorized and why it did not capture the full speedup, use GCC's diagnostics: -fopt-info-vec-optimized (what got vectorized) and -fopt-info-vec-missed (why something did not) — the fastest way to read the compiler's mind.

Hand-written intrinsics: AVX2 in practice ​

When auto-vectorization cannot get there (a non-standard access pattern, a specific instruction, or you simply must squeeze out every last drop), write intrinsics by hand. Intrinsics are compiler-builtin C functions that map to individual SIMD instructions, in the header <immintrin.h>. The heart of the AVX2 dot product:

C++
#include <immintrin.h>
float dot_avx2(const float* a, const float* b) {
    __m256 v0 = _mm256_setzero_ps(), v1 = _mm256_setzero_ps(),
           v2 = _mm256_setzero_ps(), v3 = _mm256_setzero_ps();
    for (int i = 0; i < N; i += 32) {           // 4 vectors × 8 floats = 32
        v0 = _mm256_fmadd_ps(_mm256_loadu_ps(a+i),    _mm256_loadu_ps(b+i),    v0);
        v1 = _mm256_fmadd_ps(_mm256_loadu_ps(a+i+8),  _mm256_loadu_ps(b+i+8),  v1);
        v2 = _mm256_fmadd_ps(_mm256_loadu_ps(a+i+16), _mm256_loadu_ps(b+i+16), v2);
        v3 = _mm256_fmadd_ps(_mm256_loadu_ps(a+i+24), _mm256_loadu_ps(b+i+24), v3);
    }
    // horizontally merge the 4 accumulators
    __m256 s = _mm256_add_ps(_mm256_add_ps(v0, v1), _mm256_add_ps(v2, v3));
    alignas(32) float tmp[8]; _mm256_store_ps(tmp, s);
    float acc = 0; for (int i = 0; i < 8; ++i) acc += tmp[i];
    return acc;
}

A few key points:

  • __m256 is the 256-bit vector type, holding 8 floats.
  • _mm256_loadu_ps: loads 8 floats from memory (u = unaligned, misaligned addresses allowed; the aligned variant _mm256_load_ps requires the data to be 32-byte aligned).
  • _mm256_fmadd_ps(a, b, c): a*b + c — FMA does the multiply and the add in one instruction (AVX2 + FMA only; compile with -mfma).
  • 4 accumulators v0..v3: break the dependency chain (04-02) and let 4 FMAs run in parallel, filling the execution ports.
  • Horizontal merge: the SIMD accumulators eventually have to be reduced to one scalar, and this step is heavy on scalar operations, so do not open too many accumulators either (with 8, the payoff is already worse than with 4).

This code fully exploits SIMD width + FMA + multiple accumulators, and measured 20×. But the price is poor portability (hard-bound to AVX2+FMA; older CPUs cannot run it) and poor readability. So the boundary for hand-written intrinsics: if auto-vectorization can do it, do not hand-write; hand-write only when "auto cannot + it is a confirmed hotspot".

AVX-512: watch out for downclocking ​

AVX-512 sounds even fiercer (512-bit, 16 floats at a time), but there is a trap: on some Intel CPUs, executing AVX-512 instructions triggers license-based frequency downclocking — the CPU judges AVX-512's power/thermal demands to be high and proactively drops the all-core frequency by a notch, and the downclocking persists for a while even after the AVX-512 instructions end (waiting to "cool down"). The result: a small stretch of AVX-512 code can slow down the other parts of the entire program — a bad trade.

This trap has been improving across Intel generations (from Ice Lake on, lightweight AVX-512 does not downclock; the Skylake-X generation had it worst), and AMD only introduced AVX-512 with Zen 4, under a different policy. In practice, SSE/AVX2 + FMA is the safe sweet spot: the gains are already large and there is no downclocking trap; AVX-512 should be used either on newer CPUs confirmed not to downclock, or with the "lightweight" instruction subset — and you must benchmark. The local 5800H here is Zen 3, supporting only up to AVX2 (no AVX-512), which is why this volume's SIMD examples are all AVX2-level.

CPU dispatch: function multiversioning ​

Hand-written intrinsics are hard-bound to one instruction set, so portability is poor. The fix is CPU dispatch / function multiversioning: write the same function in several versions (an SSE version, an AVX2 version, a scalar version) and pick at runtime by CPU capability. GCC's builtin __builtin_cpu_supports("avx2") queries capability; the more elegant route is __attribute__((target_clones("avx2","default"))), which has the compiler generate the multiple versions plus the dispatch automatically:

C++
__attribute__((target_clones("avx2","default")))
float dot(const float* a, const float* b) { /* one generic body; the compiler compiles it once per target */ }

The compiler generates an AVX2 version and a default version, and at the first call it detects the CPU and picks one (the IFUNC mechanism). This way one copy of source runs its own optimized version on each different CPU. The costs: a bigger binary (multi-version code) and a tiny dispatch overhead on the first call.

Outlook: std::simd (C++26) ​

C++26 standardizes std::simd (in <experimental/simd>), providing a portable SIMD abstraction: you write "process this with width-N SIMD" and the compiler maps it to each platform's concrete instructions. This hands the portability problem of hand-written intrinsics to the standard library. vol6 only mentions this as an outlook here; the deep dive waits for C++26 to spread.

Compress this article into a few sentences: SIMD gives a theoretical ×4/×8/×16, and the measured hand-written AVX2 dot product landed at ~20× (SIMD width + FMA + multiple accumulators stacked); the conditions for auto-vectorization are an estimable iteration count, no aliasing, and no complex branches; modern GCC (GCC 10+) vectorizes FP reductions in ordered mode by default (order preserved, bit-exact results), but a single accumulator pins the dependency chain shut — capturing the full SIMD width takes multiple accumulators (hand-written, or -ffast-math so the compiler dares to split the lanes), it is not just "turn on -O3/-ffast-math and done"; read the diagnostics with -fopt-info-vec; hand-written intrinsics squeeze SIMD dry but port poorly — use them only on hotspots where auto-vectorization cannot do the job; AVX-512 has the downclocking trap, SSE/AVX2+FMA is the safe sweet spot, and you must benchmark; CPU dispatch (target_clones / __builtin_cpu_supports) solves portability, and std::simd (C++26) is the future.

References ​

  • Agner Fog, Optimizing assembly, §5 "Vector programming" + §13 "AVX-512". Local copy.
  • GCC docs for -ftree-vectorize / -fopt-info-vec / target_clones / __builtin_cpu_supports
  • Intel Intrinsics Guide (intel.com/intrinsics-guide) — the intrinsics quick reference
  • This article's measurement code: code/volumn_codes/vol6-performance/ch04/simd.cpp

pdf-latest-4-g85128cc · 85128cc · 2026-10-05