Compiler optimization boundaries and binary size
This is vol6's last chapter, closing things out from the compiler and linker perspective. The earlier chapters covered how to write "C++ code that runs fast on the hardware"; this one covers what the compiler can do for you and what it can't, and how to work with it (give it line of sight, stay out of its way).
Four articles:
- 07-01 -O levels and blockers:
-O2is the release sweet spot (measured: -O0→-O2 is 4× faster), -O3 is occasionally slower than -O2 (honest); three kinds of blockers (cross-TU / aliasing / volatile). - 07-02 LTO and PGO: LTO cross-TU inlining measured at 3.9×; PGO shows no gain on microbenchmarks (an honest null — that one-time 4× was instrumentation overhead), its value is on large codebases.
- 07-03 Linking, multi-compiler, metaprogramming: the PIC cost of dynamic linking (CSAPP ch7), the small gap between GCC/Clang/MSVC, and the size side of compile-time metaprogramming.
- 07-04 Size optimization:
-Os/--gc-sections/ template bloat control; size ↔ speed is often a reverse tradeoff.
The spirit running through this chapter matches ch04: the compiler is your performance teammate, and your job is to "stay out of its way" — let it see the implementations (LTO), trust your no-aliasing promise (__restrict), and don't use volatile to disable its optimizations. Turning on LTO + --gc-sections by default in release is a free lunch; PGO is the extra course on large projects.
Local measurements cover 07-01/02/04; 07-03 is mostly conceptual (linking mechanism follows CSAPP ch7, multi-compiler comparison follows Agner vol. 1 §8), so we don't reinvent those experiments.