Skip to content
CppCon2025

Cache-Friendly C++

Jonathan Müller

These are notes from Jonathan Müller's CppCon 2025 talk. Jonathan has spent years in low-latency C++ (at think-Cell when the talk was given, later at LSEG working on HFT market-data feeds), and this time he opens with a counter-intuitive phenomenon that sends quite a few blood pressures soaring: std::unordered_set lookup is O(1), std::vector linear search is O(n), yet run them on a real machine and, at small data volumes, vector ends up crushing unordered_set. There is only one direction the answer can point in: the CPU cache.

This talk is a companion piece to the microbenchmark one. Part three of that series already pointed at the branch predictor and the cache conspiring to cheat, but never expanded on the cache itself. This talk takes the cache apart from the roots: why it exists, what unit it moves data in, how much latency differs level to level, and how every decision you make in C++ — which container you pick, which data type, how you lay out a struct — lands on cache behavior.

The notes are split into four parts, following the arc of "first break the complexity superstition, then build cache intuition, and finally land it in code." All experiments were run on the same machine: Arch Linux / WSL2, AMD Ryzen 7 9700X (Zen 5), GCC 16.1.1, -std=c++20, with a cache hierarchy of L1d 48 KiB per core, L2 1 MiB per core, and a shared 32 MiB L3. The numbers you get will be different, but the direction of the conclusions is the same.

Contents ​

pdf-latest-4-g85128cc · 85128cc · 2026-10-05