Skip to content

Measurement pitfalls and environment readiness: a 16-item checklist ​

Why we need a checklist ​

ch01-01 covered the third way a microbenchmark can lie to you: system noise drowning out the signal. That article answered "why is there noise"; this one answers "how exactly do I turn it off". The 16 items below are the environmental traps you are most likely to step on when running credible microbenchmarks on Linux. Each follows the "pitfall → why it distorts → how to avoid it" pattern, and most come with a command you can copy straight into your terminal.

But before we dive in, let's repeat the single most important boundary from ch01-01: these noise-elimination techniques should only be used when you are making a relative A/B comparison. If what you are evaluating is "how fast does this actually feel to the user", you should instead replicate the real environment (keep the noise, keep the DFS, keep the neighboring processes) and then handle it with statistical methods — that is the job of production measurement in ch01-05. This page serves the microbenchmark scenario of "I want to compare two implementations cleanly".

The 16 pitfalls ​

To make them easier to remember, they are grouped into four buckets by nature.

Group 1: Frequency and power (the biggest source of swings in your numbers) ​

#PitfallWhy it distortsHow to avoid
1CPU frequency scaling (DVFS)governor=ondemand leaves the frequency floating freely, and GBench even prints ***WARNING*** CPU scaling enabled at startupsudo cpupower frequency-set -g performance to lock the top frequency
2Turbo BoostA single core bursts to a high frequency; cold start differs from steady state, and the boost drops once the part heats upDisable Turbo in the BIOS; or lock the frequency; when measuring steady state, warm it up thoroughly first (this is exactly what warmup is about)

If you leave these two unsolved, two runs of the same code differing by 10% is perfectly normal. It is especially severe on laptops (limited cooling, so Turbo keeps entering and exiting).

Group 2: Caches, memory, and address translation ​

#PitfallWhy it distortsHow to avoid
3Cold start vs steady stateThe first access misses the cache (it goes to DRAM) while later accesses hit — a 10–100x gapThe framework's estimation phase has already warmed things up; if you want to measure the cold path, drop the pages with posix_fadvise(fd, POSIX_FADV_DONTNEED)
4Page faultsThe first touch of a page triggers a soft fault (microsecond-scale), inflating a single operation by tens of timesLock the pages with mlockall(MCL_CURRENT | MCL_FUTURE); or touch every page once in advance
9NUMAOn multi-socket machines a cross-node memory access costs 2–4x more, so what you measured as "memory bandwidth" quietly becomes "interconnect bandwidth"numactl --cpunodebind=0 --membind=0 ./bench to bind threads and memory to the same node
15ASLR / code layoutA different PIE base address makes instruction-cache (icache) and branch-predictor alignment jitter by 10–20%; it also drives "memory layout bias" (Mytkowicz 2009)Add -no-pie for fine-grained microarchitectural measurements; to cancel layout bias, use random interleaving

Group 3: Scheduling and interference ​

#PitfallWhy it distortsHow to avoid
5Context switches / interruptsYour code gets scheduled away, producing long-tail outlier samplesPin to one core with taskset -c <core>; report the median (not the mean)
8CPU pinningThreads migrate across cores and the cache goes cold every timetaskset -c 3 ./bench (pick one core and do not let the OS wander around)
10SMT / hyperthread contentionThe other thread on the same physical core eats your execution unitsDisable hyperthreading in the BIOS; or pin to physical cores only with taskset (use one of every two sibling cores)
11Timer resolutionclock() on nanosecond-scale work is pure noise (not enough resolution)std::chrono::steady_clock (see ch00-02); or read cycles from perf stat

Group 4: Tool usage and statistics ​

#PitfallWhy it distortsHow to avoid
6Dead code optimized awayThe result is never consumed → DCE deletes the loop (see foo() in ch01-01)DoNotOptimize / doNotOptimizeAway; note that it does not stop the expression itself from being folded away (ch01-02)
7No release build plus debug infoPerformance numbers from -O0 are meaningless; plain -O2 without -g cannot be source-annotatedStandardize on RelWithDebInfo (-O2 -g); for profiling, add -fno-omit-frame-pointer (otherwise stacks break and the flame graph explodes)
12Mean vs medianMicrobenchmark results are right-skewed (long tail), dragging the mean upwardReport the median + IQR; GBench Repetitions + ReportAggregatesOnly (ch01-02)
13Too few samplesConfidence intervals so wide that A and B are indistinguishable≥30 samples; report the 95% CI; judge A/B significance with the Mann-Whitney U test (ch01-04)
14Unstable across repeated runsThe environment is not pinned down, so results drift between runsTake ≥3 runs and keep the steadiest; perf stat -r 5 does the repetition for you
16PEBS skidA sampled event "skids" a few instructions before it lands on the instruction actually responsibleUse events with the :pp / :ppp (precise IP) suffix, such as MEM_LOAD_RETIRED.L3_MISS:ppp

Sixteen items looks like a lot, but the core is one sentence: control everything controllable (frequency, cores, memory layout), really consume every result that should be consumed (DoNotOptimize), and treat the numbers as a distribution (median + multiple repeated rounds). The rest depends on the scenario: for micro work, do as many as you can; for production evaluation, do none of them (replicate reality instead).

One-shot health check: perf-env-check.sh ​

Checking all of these by hand before every measurement session is tedious, so we compressed it into one script. It checks only, never modifies (actions that require sudo, such as changing the governor or disabling Turbo, are left to your own judgment) and prints out whatever problems it finds:

Expand codeCollapse34 lines
bash
#!/usr/bin/env bash
# perf-env-check.sh — environment health check for credible microbenchmarks (check only, never modify)
set -u

ok()   { printf "  ✓ %s\n" "$1"; }
warn() { printf "  ⚠ %s\n" "$1"; }

echo "=== CPU governor(应=performance)==="
g=$(cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor 2>/dev/null)
[ "$g" = performance ] && ok "governor=performance" || warn "governor=$g(DVFS 会浮动)。修:sudo cpupower frequency-set -g performance"

echo "=== Turbo Boost(Intel)==="
if [ -f /sys/devices/system/cpu/intel_pstate/no_turbo ]; then
  nt=$(cat /sys/devices/system/cpu/intel_pstate/no_turbo)
  [ "$nt" = 1 ] && ok "Turbo 已关" || warn "Turbo 开着(no_turbo=$nt),冷热启动数字会差"
else
  echo "  · 非 intel_pstate 或无该接口,跳过(可在 BIOS 设)"
fi

echo "=== perf_event_paranoid(<=1 才好采样)==="
p=$(cat /proc/sys/kernel/perf_event_paranoid 2>/dev/null)
[ "${p:-3}" -le 1 ] && ok "perf_event_paranoid=$p" || warn "=$p(perf 受限)。修:sudo sysctl -w kernel.perf_event_paranoid=1"

echo "=== NUMA 拓扑(多 socket 才在意)==="
command -v numactl >/dev/null && numactl --hardware 2>/dev/null | grep -E "^available|node [0-9]+ cpus" | head -4 || warn "无 numactl"

echo "=== CPU 亲和性(应明确绑一个核,别让 OS 晃)==="
cpu=$(grep Cpus_allowed_list /proc/self/status 2>/dev/null | awk '{print $2}')
n=$(nproc 2>/dev/null)
echo "  Cpus_allowed_list=$cpu (nproc=$n) → 没绑核就 taskset -c <某个核> ./bench(别挑 0 号核,常被系统中断占用)"

echo "=== ASLR(微架构精细测时应关)==="
aslr=$(cat /proc/sys/kernel/randomize_va_space 2>/dev/null)
echo "  randomize_va_space=$aslr(2=全开;精细 icache/分支测时可 sudo sysctl -w kernel.randomize_va_space=0)"

Save it as perf-env-check.sh; one run of bash perf-env-check.sh tells you what your environment is still missing. The complete script also lives under code/volumn_codes/vol6-performance/ch01/.

Which measures fit which scenario ​

ScenarioDo these (eliminate noise)What not to do
Microbenchmark A/B comparison1/2/4/5/8/9/10/15 — do as many as possible; what you want is a clean signal-to-noise ratioDo not extrapolate the conclusions straight to production
Evaluating production performanceDo almost none of them — replicate the real environment (keep DFS, neighbors, ASLR)Do not turn off noise sources, or you are no longer measuring what users will actually experience
Profiling to find hotspots7 (-fno-omit-frame-pointer), 16 (:pp)Hotspot hunting is by definition sampling under real load

This table is where the ch01-01 line "the very hand that makes a micro clean is the hand that makes it lie to you" becomes concrete: the same set of techniques is the antidote in a microbenchmark scenario and poison in a production scenario. Getting that dosage right matters more than memorizing 16 commands.

References ​

  • easyperf.net: How to get consistent results when benchmarking on Linux (one of the direct sources of this checklist)
  • Brendan Gregg: Linux Performance (perf / task placement / NUMA)
  • Bakhvalov, D., Performance Analysis and Tuning on Modern CPUs, §2.1 Noise In Modern Systems
  • This volume's ch01-01 (the taxonomy of noise) and ch01-02 (DoNotOptimize / Repetitions / UseRealTime)

pdf-latest-4-g85128cc · 85128cc · 2026-10-05