Skip to content

Linking performance, multi-compiler comparison, and compile-time metaprogramming ​

The first two ch07 articles covered -O levels and LTO/PGO. This one closes out the three remaining compile/link-related topics: the runtime cost of dynamic linking, optimization differences among the mainstream compilers, and the performance face of compile-time metaprogramming. All three are "boundary topics" — not the lead actors in routine performance work, but they pop up in large projects, embedded systems, and extreme-optimization scenarios.

No local measurements in this one. The mechanics of dynamic linking follow CSAPP ch7, the multi-compiler comparison follows Agner vol1 §8, and compile-time metaprogramming follows Agner vol1 §15 — we only cover the performance side and don't reinvent the experiments.

Dynamic linking and PIC: the cost of position independence ​

CSAPP Chapter 7, Linking, covers static vs dynamic linking. From a performance angle, dynamic linking (.so/.dylib/.dll) carries several runtime costs:

  • Symbol resolution: the first time you call a function in a dynamic library, the runtime linker (ld.so) has to search the symbol table and bind the address (lazy binding), or bind everything up front at startup (now binding). This is a first-call latency.
  • PIC (position-independent code): dynamic-library code must be loadable at any address, so it goes through GOT (Global Offset Table) indirect addressing — every access to a global variable or external function adds one more GOT lookup. This is a constant per-access cost (a few extra instructions).
  • PLT (Procedure Linkage Table): calls to external functions go through a PLT indirect jump, one extra hop compared with a direct call.
  • icache pressure: dynamic-library code is scattered across different load addresses, so cross-library calls add icache misses.

Each of these costs is tiny per call (nanoseconds), but it shows up in "extremely high-frequency cross-library calls + very short functions" territory (say, a tight loop calling a small function in another .so). A few counters:

  • Inline away the hot path's cross-library calls (LTO works within a library, not across libraries; or move the hot function into the main program).
  • -Wl,-z,now (bind everything up front at startup, so the first-call latency doesn't hit mid-run; a good fit for long-running services).
  • Statically link the hot-path libraries (trading upgrade independence for performance).

CSAPP ch7 has the full walkthrough of the mechanics; vol6 only covers the performance side. For most applications these costs are negligible — it's real-time games, HFT, and embedded that fine-tune them.

Multi-compiler comparison: GCC vs Clang vs MSVC ​

The three mainstream compilers' optimization capability is broadly on par, differing in the details. Agner Fog compared them in vol1 §8; a few observations (these shift between versions — check the latest):

  • Optimization strength: at -O2/-O3, the overall performance spread across the three is usually within single-digit percentage points, and which one wins depends on the specific code.
  • Vectorization: GCC has historically vectorized aggressively; Clang/LLVM's vectorization framework (led by Nadav Rotem) is also strong and sometimes smarter; MSVC's auto-vectorization is comparatively conservative (though hand-written intrinsics are all the same).
  • Code generation: Clang has the best error messages and the most accurate diagnostics; GCC has the broadest platform support; MSVC is the de facto standard under the Windows ABI.
  • Cross-platform: GCC/Clang span Linux/macOS/Windows (MinGW/clang-cl); MSVC is mostly Windows.

The practical advice: don't switch toolchains because "I heard compiler X is faster" — the gap is small, and it shifts between versions. Pick a compiler on platform support + team familiarity + diagnostic quality; the performance gap is not the deciding factor. If you really want to squeeze out the last drop, PGO + LTO is far more effective than switching compilers. Just run benchmarks periodically to confirm your compiler choice has no obvious deficit.

Note: different compilers have incompatible ABIs (Itanium ABI vs MSVC ABI), so mixing them (say, handing a GCC-built library to MSVC) needs an extern "C" interface as isolation. That's an engineering problem, not a performance one.

Compile-time metaprogramming: the performance face of templates and constexpr ​

C++ templates and constexpr can push computation to compile time: compilation finishes, the computation is done, and runtime cost is zero. This is "true zero-overhead abstraction" (at runtime). But the cost has two sides.

Runtime: zero or close to it ​

  • constexpr/consteval functions: evaluation finishes at compile time, and the runtime result is just a constant. For example, constexpr int fib(int n) computing fib(10) at compile time means runtime sees just the literal 55 — zero runtime cost.
  • Template computation: template<int N> struct Fact { static constexpr int v = N * Fact<N-1>::v; };, computed at compile time.
  • Type computation (the dispatch of std::tuple, std::variant) is generated at compile time; at runtime it is direct, optimized code (often inlined away).

So "whatever can be computed at compile time, compute at compile time" is one of C++'s performance principles: moving compute from runtime to compile time is free.

The cost: compile time and size ​

  • Compile time: template instantiation is one of the compiler's heaviest jobs. Heavy-template C++ projects routinely take tens of seconds to minutes per build. This is a long-standing pain point in the C++ world (modules, since C++20, ease it).
  • Binary size: templates generate one copy of the code per type (vector<int>, vector<double>, vector<string> are three copies) — template bloat. ch07-04 covers the counters (extern template, factoring out shared logic).
  • Readability/error messages: heavy-template code is famous for how unreadable its error messages are.

The practical tradeoff: for small computations on the hot path, push them to compile time with constexpr (constexpr int kTable[N] = ...); don't over-templatize just for the sake of "compile time" (template bloat + compile time + readability cost). constexpr is more restrained and more modern than "template metaprogramming" — reach for constexpr/consteval first.

References ​

  • CSAPP Chapter 7, Linking — the mechanics of static/dynamic linking, GOT/PLT, symbol resolution
  • Agner Fog, Optimizing software in C++, §8 "Different C++ compilers" (compiler comparison) + §15 "Metaprogramming" (the size side of compile-time metaprogramming). Local copy
  • GCC/Clang/MSVC documentation for their respective -O behavior, -fpic/-fPIC, and constexpr support
  • ch07-04 size optimization (this volume, counters to template bloat)

Three closing takeaways: dynamic linking has runtime cost (PIC indirection, symbol resolution, icache) — tiny per call, visible only under extremely high-frequency short cross-library calls, mechanics in CSAPP ch7; the three big compilers' performance gap is small (single-digit percentage points), don't switch toolchains for "faster", PGO+LTO is more effective; compile-time metaprogramming is zero-cost at runtime, but the bill is compile time + template size — prefer constexpr/consteval over heavy templates. In routine optimization these three are supporting cast, but in large-project build engineering, embedded, and extreme optimization they take the lead.

pdf-latest-4-g85128cc · 85128cc · 2026-10-05