Skip to content

Multicore performance ​

vol5 covered the correctness of synchronization primitives, memory ordering, and lock-free data structures thoroughly. This chapter doesn't repeat any of that machinery; it answers one performance question: how to measure and fix the performance decay multicore brings, and how many nanoseconds each synchronization style costs.

Three articles:

  • 05-01 False sharing: two cores frequently writing different variables sitting on the same 64-byte cacheline trigger MESI coherence invalidate round-trips, measured an order of magnitude slower (about 18x in a single run, and the multiplier swings a lot between runs). The fix is alignas(64).
  • 05-02 NUMA and the scalability curve: on multi-socket machines, cross-node memory access latency is 2-4x higher; the scalability curve (measured 1→8 threads at a sublinear 2.53x) diagnoses "how much performance more cores actually buy you"; Amdahl vs Gustafson; thread affinity (core pinning); thread creation and stack cost.
  • 05-03 Lock vs lock-free cost: an uncontended mutex is nanosecond-level (about 3.6x an atomic), but the cost explodes under contention; lock-free is not a silver bullet (ABA, retry storms, memory reclamation), and sharded locks routinely beat lock-free.

Boundary: "how to write correct synchronization, atomic-operation memory ordering, and lock-free implementations" belongs to vol5; vol6 only answers "how much each costs and which to pick in which scenario."

This machine is WSL2 on a single socket with a single NUMA node, so the cross-node NUMA penalty can't be measured here (05-02 marks this honestly; that content is drawn from multi-socket server practice). False sharing, scalability, and lock overhead were all measured on this machine.

In this chapter ​

pdf-latest-4-g85128cc · 85128cc · 2026-10-05