Skip to content

std::function's small buffer optimization: the cost of type erasure ​

The convenience and the cost of type erasure ​

std::function is C++'s most convenient "hold any callable" container: function pointers, lambdas, function objects, bind expressions — as long as the signature matches, anything goes in. Its implementation relies on type erasure: instead of baking the callable's type into a template parameter, it calls through a uniform internal interface (usually a virtual function or a table of function pointers).

That convenience has a cost. Let's measure it (on this machine):

text
===== std::function SBO =====
调用:
  调用 - 函数指针:               0.26 ns
  调用 - function+小lambda(SBO): 1.61 ns
  调用 - 直接 lambda(对照):     0.25 ns

构造 1000000 次:
  function 装函数指针:     2.0 ns/次
  function 装小lambda(SBO): 2.3 ns/次
  function 装大lambda(堆分配): 19.6 ns/次 ← 堆分配开销

sizeof(std::function<int(int)>) = 32

Two costs:

1. The call is indirect, about 6x slower than a direct lambda. A direct lambda (0.25 ns) is as fast as a function pointer (0.26 ns) — both are direct calls the compiler can inline; std::function (1.61 ns) has to go through the type-erased indirect call (vtable/function-pointer lookup plus a jump), about 6x. Same story as the virtual functions in 06-01: indirect calls block inlining.

2. Construction may heap-allocate. When std::function holds a callable, it has to store that callable's state. Most implementations have SBO (Small Buffer Optimization): a small buffer is reserved inside the std::function object itself (on this machine, libstdc++'s std::function<int(int)> is 32 bytes). Captures with small state (≤ the SBO threshold, usually 16-24 bytes) are stored inline, no heap allocation; captures with large state (over the threshold) have no option but to new a block on the heap.

Measured construction cost: a small lambda (SBO hit) takes 2.3 ns per construction, a large lambda (over SBO, heap-allocated) takes 19.6 ns, 8.5x. That gap mostly comes from the cost of a single new/delete heap allocation (on the order of tens of nanoseconds).

When these two costs bite ​

The 6x call cost: for a callback invoked "once in a while", it doesn't matter (the callback isn't on the hot path); for a callback invoked "a million times per frame", 6x is real money. Take an event dispatcher: if every dispatch goes through std::function, the call overhead can become the bottleneck, and switching to a template (compile-time polymorphism) or a function pointer is much faster.

The construction heap-allocation cost: this is the easier trap to fall into. Consider code like this:

C++
// Hot path: repeatedly constructing a function with a big capture
for (auto& item : items) {
    std::function<void(int)> f = [item, ctx](int x) { /* big capture */ };
    dispatch(f);
}

Every loop iteration constructs a std::function, and if the capture is big (over SBO), every iteration does a new/delete: heap allocation + cache misses + possible malloc lock contention (under multithreading). This pattern of "repeatedly constructing a function on the hot path" is a performance black hole. A few common fixes:

  • Use a template parameter (compile-time polymorphism): make the callback type a template parameter, eliminating type erasure. The cost is that call sites must know the type at compile time.
  • A fixed-signature function pointer: if the callback captures nothing, just use void(*)(int) — zero overhead.
  • Reuse the std::function object: construct it once outside the loop, and only mutate its state inside the loop (though mutating state may still heap-allocate).
  • Avoid unnecessary captures: the less a lambda captures, the more likely it is to hit SBO.

SBO is the same idea as string's SSO ​

SBO is the same idea as std::string's SSO (Small String Optimization): both keep a small buffer inside the object — small stays inline, only big goes to the heap. Both resolve the tension between "type erasure / dynamic size" and "avoiding heap allocation on the hot path". The mechanics of SBO/SSO (why the threshold is 16-24 bytes, how it cooperates with the ABI) belong to vol3/vol4; vol6 only covers the layer of "it affects the heap-allocation cost of hot-path construction".

The sizeof of std::function varies by implementation (libstdc++ 32 bytes, libc++ 48 bytes, MSVC different again), and the SBO threshold varies with it. So "will this lambda of mine trigger a heap allocation" is a question for sizeof or for reading the implementation — but the general advice is: don't depend on hitting SBO on the hot path; big captures belong in templates.

One sentence to wrap up: std::function has two costs — the call is indirect (about 6x slower than a direct lambda), and construction may heap-allocate (triggered by a big capture, about 8.5x more expensive than SBO); SBO lets small captures (≤16-24B) live inside the object with no heap allocation, while big captures heap-allocate; on the hot path, avoid repeatedly constructing std::function with a big capture — that is a heap-allocation black hole, and the fixes are a template parameter, a function pointer, reusing the object, or cutting captures; SBO shares its idea with string's SSO, and the mechanics belong to vol3/vol4.

References ​

  • cppreference std::function — type-erasure semantics, SBO notes
  • Sutter/Stroustrup CppCoreGuidelines F.50 — when to use function vs template vs function pointer
  • Agner Fog, Optimizing software in C++, object/container overhead. Local copy
  • The measurement code for this article: code/volumn_codes/vol6-performance/ch06/function_sbo.cpp

pdf-latest-4-g85128cc · 85128cc · 2026-10-05