Linking performance, multi-compiler comparison, and compile-time metaprogramming
The first two ch07 articles covered -O levels and LTO/PGO. This one closes out the three remaining compile/link-related topics: the runtime cost of dynamic linking, optimization differences among the mainstream compilers, and the performance face of compile-time metaprogramming. All three are "boundary topics" — not the lead actors in routine performance work, but they pop up in large projects, embedded systems, and extreme-optimization scenarios.
No local measurements in this one. The mechanics of dynamic linking follow CSAPP ch7, the multi-compiler comparison follows Agner vol1 §8, and compile-time metaprogramming follows Agner vol1 §15 — we only cover the performance side and don't reinvent the experiments.
Dynamic linking and PIC: the cost of position independence
CSAPP Chapter 7, Linking, covers static vs dynamic linking. From a performance angle, dynamic linking (.so/.dylib/.dll) carries several runtime costs:
- Symbol resolution: the first time you call a function in a dynamic library, the runtime linker (
ld.so) has to search the symbol table and bind the address (lazy binding), or bind everything up front at startup (now binding). This is a first-call latency. - PIC (position-independent code): dynamic-library code must be loadable at any address, so it goes through GOT (Global Offset Table) indirect addressing — every access to a global variable or external function adds one more GOT lookup. This is a constant per-access cost (a few extra instructions).
- PLT (Procedure Linkage Table): calls to external functions go through a PLT indirect jump, one extra hop compared with a direct call.
- icache pressure: dynamic-library code is scattered across different load addresses, so cross-library calls add icache misses.
Each of these costs is tiny per call (nanoseconds), but it shows up in "extremely high-frequency cross-library calls + very short functions" territory (say, a tight loop calling a small function in another .so). A few counters:
- Inline away the hot path's cross-library calls (LTO works within a library, not across libraries; or move the hot function into the main program).
-Wl,-z,now(bind everything up front at startup, so the first-call latency doesn't hit mid-run; a good fit for long-running services).- Statically link the hot-path libraries (trading upgrade independence for performance).
CSAPP ch7 has the full walkthrough of the mechanics; vol6 only covers the performance side. For most applications these costs are negligible — it's real-time games, HFT, and embedded that fine-tune them.
Multi-compiler comparison: GCC vs Clang vs MSVC
The three mainstream compilers' optimization capability is broadly on par, differing in the details. Agner Fog compared them in vol1 §8; a few observations (these shift between versions — check the latest):
- Optimization strength: at
-O2/-O3, the overall performance spread across the three is usually within single-digit percentage points, and which one wins depends on the specific code. - Vectorization: GCC has historically vectorized aggressively; Clang/LLVM's vectorization framework (led by Nadav Rotem) is also strong and sometimes smarter; MSVC's auto-vectorization is comparatively conservative (though hand-written intrinsics are all the same).
- Code generation: Clang has the best error messages and the most accurate diagnostics; GCC has the broadest platform support; MSVC is the de facto standard under the Windows ABI.
- Cross-platform: GCC/Clang span Linux/macOS/Windows (MinGW/clang-cl); MSVC is mostly Windows.
The practical advice: don't switch toolchains because "I heard compiler X is faster" — the gap is small, and it shifts between versions. Pick a compiler on platform support + team familiarity + diagnostic quality; the performance gap is not the deciding factor. If you really want to squeeze out the last drop, PGO + LTO is far more effective than switching compilers. Just run benchmarks periodically to confirm your compiler choice has no obvious deficit.
Note: different compilers have incompatible ABIs (Itanium ABI vs MSVC ABI), so mixing them (say, handing a GCC-built library to MSVC) needs an
extern "C"interface as isolation. That's an engineering problem, not a performance one.
Compile-time metaprogramming: the performance face of templates and constexpr
C++ templates and constexpr can push computation to compile time: compilation finishes, the computation is done, and runtime cost is zero. This is "true zero-overhead abstraction" (at runtime). But the cost has two sides.
Runtime: zero or close to it
constexpr/constevalfunctions: evaluation finishes at compile time, and the runtime result is just a constant. For example,constexpr int fib(int n)computingfib(10)at compile time means runtime sees just the literal55— zero runtime cost.- Template computation:
template<int N> struct Fact { static constexpr int v = N * Fact<N-1>::v; };, computed at compile time. - Type computation (the dispatch of
std::tuple,std::variant) is generated at compile time; at runtime it is direct, optimized code (often inlined away).
So "whatever can be computed at compile time, compute at compile time" is one of C++'s performance principles: moving compute from runtime to compile time is free.
The cost: compile time and size
- Compile time: template instantiation is one of the compiler's heaviest jobs. Heavy-template C++ projects routinely take tens of seconds to minutes per build. This is a long-standing pain point in the C++ world (modules, since C++20, ease it).
- Binary size: templates generate one copy of the code per type (
vector<int>,vector<double>,vector<string>are three copies) — template bloat. ch07-04 covers the counters (extern template, factoring out shared logic). - Readability/error messages: heavy-template code is famous for how unreadable its error messages are.
The practical tradeoff: for small computations on the hot path, push them to compile time with constexpr (constexpr int kTable[N] = ...); don't over-templatize just for the sake of "compile time" (template bloat + compile time + readability cost). constexpr is more restrained and more modern than "template metaprogramming" — reach for constexpr/consteval first.
References
- CSAPP Chapter 7, Linking — the mechanics of static/dynamic linking, GOT/PLT, symbol resolution
- Agner Fog, Optimizing software in C++, §8 "Different C++ compilers" (compiler comparison) + §15 "Metaprogramming" (the size side of compile-time metaprogramming). Local copy
- GCC/Clang/MSVC documentation for their respective
-Obehavior,-fpic/-fPIC, andconstexprsupport - ch07-04 size optimization (this volume, counters to template bloat)
Three closing takeaways: dynamic linking has runtime cost (PIC indirection, symbol resolution, icache) — tiny per call, visible only under extremely high-frequency short cross-library calls, mechanics in CSAPP ch7; the three big compilers' performance gap is small (single-digit percentage points), don't switch toolchains for "faster", PGO+LTO is more effective; compile-time metaprogramming is zero-cost at runtime, but the bill is compile time + template size — prefer constexpr/consteval over heavy templates. In routine optimization these three are supporting cast, but in large-project build engineering, embedded, and extreme optimization they take the lead.