Skip to content
CppCon2025

Why 99% of C++ Microbenchmarks Lie – and How to Write the 1% that Matter!

Kris Jusiak

These are notes from Kris Jusiak's CppCon 2025 talk. Kris is the author of [Boost].UT and has spent years tinkering with compile-time computation and testing frameworks. This time he locks onto a question that sends your blood pressure through the roof: you write a benchmark, it prints a beautiful nanosecond number, you optimize against it, merge the code, ship — and in production nothing moves, or things even get slower. Kris's answer stings: that benchmark of yours is, with high probability, lying — and not in just one place, but in several spots chained together.

The notes are split into five parts, in the order of exposing the liar layer by layer: first how the compiler optimizes the very loop you wanted to measure out of existence; then noise and bias, two kinds of error with completely different natures; next how the branch predictor and the cache conspire to hand you a fake number; then latency versus throughput, two dimensions that constantly get conflated; and finally the deadliest one of all — a faster microbenchmark and a faster whole program are simply not the same thing.

About the local environment

Every experiment in this series was run on the same machine: Arch Linux / WSL2, AMD Ryzen 7 9700X (Zen 5 architecture), GCC 16.1.1, -std=c++20. Nothing is pinned, no clocks are locked, no background is silenced — it is just a WSL2 environment running loose. That means the noise is larger than on a properly tuned setup, but it happens to make "what noise looks like" easier to see. The numbers on your machine will differ; the direction of the conclusions will not.

Contents ​

pdf-latest-4-g85128cc · 85128cc · 2026-10-05