Skip to content

std::function, std::invoke, and Callable Objects ​

While writing an event system, we ran into a very practical problem: the callbacks we had to store came in too many shapes — plain functions, member functions, lambdas, functors, all of them at once. Store a lambda with auto? Every lambda has its own distinct type, so we can't put them into the same container. Function pointers, meanwhile, can only point at a limited range of targets. That is exactly the job std::function exists for: it uses type erasure to unify all kinds of callable objects into a single type. But type erasure isn't free, so we have to ask: how big is the cost, exactly? And is there a way to get the best of both worlds?

In the first article we deferred a few details to this one: how type erasure hides every kind of closure type behind a unified interface, and where that fixed overhead goes. Along the way we will also make clear exactly what std::invoke unifies. This article takes things apart one by one; by the end, you will be able to weigh for yourself which kind of callback fits which scenario.


What Counts as a Callable in C++ ​

Let's spend a minute counting up what qualifies as a callable in C++. You can take the definition literally: anything that can be called with () syntax (being called via std::invoke counts too). Plain functions and function pointers are the most basic forms — you can call them directly, or indirectly through a function pointer. A functor is a class object with an overloaded operator(). A lambda is at heart an anonymous functor generated by the compiler, as we took apart in the earlier articles. A member function pointer points at a class's member function, and calling one has to bring an object instance along. Beyond those, there are the objects wrapped by std::function, and the results of std::bind.

So where is the trouble? The invocation syntax of these forms is not uniform. We call plain functions directly; member functions have to be written as obj.*ptr or obj->*ptr; functors and lambdas are the easy ones — f(args) and the call goes through. If you wanted to write a generic function that invokes all of them uniformly, back in the pre-C++17 days you had to stockpile a whole pile of template specializations. Once std::invoke arrived, one function settled it.


std::function: A Type-Erasing Function Wrapper ​

std::function is the general-purpose function wrapper introduced in C++11; its definition lives in the <functional> header. We can hand it any callable that matches the required signature, and it will store, copy, and invoke it. Its whole capability fits in one sentence: unify callable objects of different types into a single type.

Expand codeCollapse26 lines
C++
#include <functional>
#include <iostream>

int add(int a, int b) { return a + b; }

struct Multiplier {
    int factor;
    int operator()(int x) const { return x * factor; }
};

void demo_std_function() {
    std::function<int(int, int)> func;

    // Store a plain function
    func = add;
    std::cout << func(3, 4) << "\n";   // 7

    // Store a lambda
    func = [](int a, int b) { return a * b; };
    std::cout << func(3, 4) << "\n";   // 12

    // Store a functor
    func = Multiplier{5};
    std::cout << func(10) << "\n";     // Compile error: signature mismatch
    // Multiplier's operator() takes only one argument, but func's signature is int(int,int)
}

The Type Erasure Mechanism ​

How does std::function get away with packing things of different types into one type? Through type erasure. The name sounds intimidating, but taken apart it comes down to one thing: hide every kind of closure type behind a unified interface, so the concrete type is no longer visible from the outside. How exactly is it hidden? Read on. Internally it defines an abstract base class (the Concept) that declares a pure virtual invoke. For each concrete callable type, a derived class (the Model) is generated, and the real invoke implementation lands in that derived class. std::function itself holds nothing but a pointer to the Concept, and at call time the virtual function dispatches back to the concrete implementation.

We can simulate the whole process in code:

Expand codeCollapse49 lines
C++
#include <memory>
#include <utility>

// A simplified sketch of how std::function works
template<typename Signature>
class SimpleFunction;

template<typename R, typename... Args>
class SimpleFunction<R(Args...)> {
    // Abstract interface
    struct ICallable {
        virtual ~ICallable() = default;
        virtual R invoke(Args... args) = 0;
        virtual ICallable* clone() const = 0;
    };

    // Concrete implementation: a templated derived class stores the real callable
    template<typename T>
    struct CallableImpl : ICallable {
        T callable;
        explicit CallableImpl(T c) : callable(std::move(c)) {}

        R invoke(Args... args) override {
            return callable(std::forward<Args>(args)...);
        }

        ICallable* clone() const override {
            return new CallableImpl(callable);
        }
    };

    ICallable* ptr_ = nullptr;

public:
    SimpleFunction() = default;

    template<typename T>
    SimpleFunction(T callable)
        : ptr_(new CallableImpl<std::decay_t<T>>(std::move(callable))) {}

    SimpleFunction(const SimpleFunction& other)
        : ptr_(other.ptr_ ? other.ptr_->clone() : nullptr) {}

    ~SimpleFunction() { delete ptr_; }

    R operator()(Args... args) {
        return ptr_->invoke(std::forward<Args>(args)...);
    }
};

Set the code side by side with the description and you'll see type erasure needs just three parts: a unified abstract interface ICallable, a templated concrete implementation CallableImpl<T>, and one pointer to the interface, ptr_. When a callable goes in, its type information is "erased" — all the outside can see is an ICallable*. When the call happens, the vtable leads back to the concrete implementation.

We turned the whole store-and-invoke process into an animation. You can play it, pause it, or step through it one frame at a time: watch three kinds of callable objects get unified into one type, then dispatched back to their concrete implementations at call time:

0.0s / 50.9s
STEP 01开场

Small Object Optimization (SBO) ​

The simplified version above has one glaring flaw, and you probably spotted it at a glance: every construction does a new, so the memory lands on the heap. For a small lambda capturing only one or two ints, the heap allocation may cost more than the lambda itself. So real std::function implementations all carry small object optimization. The technique's full name is Small Buffer Optimization — literally "small buffer optimization" — so you will run into both names; the abbreviation used day to day is SBO (occasionally written SOO): a fixed-size buffer (usually 16-32 bytes) is reserved inside the std::function object, and as long as the wrapped callable is small enough, it goes straight into that buffer, and the heap allocation is skipped.

Expand codeCollapse26 lines
C++
#include <functional>
#include <iostream>
#include <array>

void demo_sbo_size() {
    // A small lambda: usually fits in the SBO buffer
    auto small = [x = 42]() { return x * 2; };
    std::function<int()> f1 = small;
    std::cout << "sizeof(std::function<int()>): "
              << sizeof(f1) << " bytes\n";
    // Usually 32-64 bytes (implementation-dependent)

    // A large lambda: exceeds the SBO buffer and triggers a heap allocation
    auto large = [data = std::array<int, 100>{}]() {
        return data.size();
    };
    std::function<std::size_t()> f2 = large;
    std::cout << "sizeof(std::function<size_t()>): "
              << sizeof(f2) << " bytes\n";
    // Same size, but a heap allocation happened internally

    // For comparison: the size of a function pointer
    std::cout << "sizeof(void(*)()): "
              << sizeof(void(*)()) << " bytes\n";
    // Usually 8 bytes (on 64-bit systems)
}

We actually put libstdc++'s SBO behavior to the test on GCC 15.2.1: std::function<int()> came in at 32 bytes. And the results were interesting: a lambda whose closure is just 4 bytes (capturing a single int) did not trigger a heap allocation, while a lambda capturing 5 ints — or one pointer — landed on the heap. GCC 15.2's SBO implementation seems fairly conservative, probably because the vtable pointer and management metadata need room of their own. libc++'s (Clang's) implementation may differ, and the exact behavior varies by version; if you want to be sure, running the test on your own compiler is the safest bet.

Verification code: the test we ran is at code/volumn_codes/vol2/ch03-lambda/test_sbo_size.cpp, on GCC 15.2.1 with -O2.

One extra reminder: SBO behavior differs a lot across compilers and versions. If you are writing performance-sensitive code, consider switching to template parameters or hand-rolled type erasure — something with predictable behavior.

We also turned where these two closures end up into an animation you can step through: how a 4-byte small closure gets tucked into the internal buffer, and why a 400-byte large one lands on the heap:

0.0s / 50.0s
STEP 01开场

Function Pointers: Zero Overhead but Limited ​

Let's set std::function aside and look back at the function pointer, handed down from the C era. It points at the code address itself — simple and efficient. Its size is one pointer (8 bytes on a 64-bit system), and a call is a single call instruction (call *%rax), with no extra layer of indirection tacked on.

Measured performance: we ran the test on GCC 15.2.1 with -O2; calling through a function pointer came out about 30% slower than a direct call (1.29x). The source of the gap is clear as well: a direct call can be fully inlined into computation instructions, while the function-pointer side still has to make an indirect call. In unoptimized code, though, both sides issue a call instruction, so the difference is smaller.

Verification code: the script we ran is at code/volumn_codes/vol2/ch03-lambda/test_function_performance.cpp.

C++
// Declaring and assigning a function pointer
int (*func_ptr)(int, int) = [](int a, int b) { return a + b; };

// Simplifying the type name with using
using BinaryOp = int(*)(int, int);
BinaryOp op = [](int a, int b) { return a + b; };
int result = op(3, 4);   // 7

The biggest limitation of function pointers is that they cannot carry context: apart from captureless lambdas, the things they can point at are plain functions and static member functions, and that's it. A lambda with any capture has no way to convert to a function pointer. The day you need to hand a this pointer or some state to a callback, the function pointer can't help you.

C++
// A captureless lambda can convert to a function pointer
int (*fp1)(int, int) = [](int a, int b) { return a + b; };  // OK

// A lambda with captures cannot convert
int x = 42;
int (*fp2)(int, int) = [x](int a, int b) { return a + b + x; };  // Compile error
FeatureFunction pointerstd::function
Size8 bytes (64-bit)32-64 bytes
Heap allocationNoneTriggered outside the SBO range
Indirection levels1 (a direct call)1 (vtable indirection)
Carries contextNoYes
Inlining friendlinessRelatively goodPoor (type erasure gets in the way)
Performance (relative to a direct call)~1.3x~7-9x

The source of the ~7-9x in the table is our run of code/volumn_codes/vol2/ch03-lambda/test_function_performance.cpp, on GCC 15.2.1 with -O2, 100 million calls in total; we will reuse these numbers in the selection guide section.


std::invoke: A Unified Invocation Interface ​

std::invoke, introduced in C++17 (defined in <functional>), takes care of a different job: unifying invocation syntax. Whether what you bring is a plain function, a member function pointer, a lambda, or a functor, we can call it with one single notation. It implements the semantics of the INVOKE expression in the C++ standard. You can treat INVOKE as the standard's name for "uniform invocation" — a name whose job is to also cover the special forms, member function pointers and member data pointers included.

Expand codeCollapse30 lines
C++
#include <functional>
#include <iostream>

struct Widget {
    void greet(const std::string& msg) {
        std::cout << "Widget says: " << msg << "\n";
    }
    int data = 42;
};

void free_func(int x) {
    std::cout << "free_func: " << x << "\n";
}

void demo_invoke() {
    Widget w;

    // A plain function
    std::invoke(free_func, 42);

    // A functor / lambda
    std::invoke([](int x) { std::cout << "lambda: " << x << "\n"; }, 99);

    // A member function pointer + object
    std::invoke(&Widget::greet, w, "hello");

    // A member data pointer + object (can be read and written)
    int val = std::invoke(&Widget::data, w);
    std::invoke(&Widget::data, w) = 100;
}

Please look at the member-function line in the code, std::invoke(&Widget::greet, w, "hello"). The traditional spelling is (w.*(&Widget::greet))("hello"), or the (w.*mem_func)("hello") form — syntax we have to look up every single time we write it. Switch to std::invoke and all you write is std::invoke(mem_func, obj, args...). Much less hassle.

How invoke Works Under the Hood ​

The implementation principle of std::invoke is not complicated; what it does is compile-time type discrimination and dispatch, and we'll walk through each case. If the object is an ordinary callable (a function pointer, a lambda, a functor), f(args...) calls it right through. When a member function pointer comes up, it picks the matching invocation syntax according to the category of the object passed in — the categories that can be passed are pointer, reference, and std::reference_wrapper. std::reference_wrapper is the standard-library type in <functional> that wraps a reference into a copyable object, precisely so that references can travel by value. When a member data pointer comes up, what it returns is a reference to the corresponding member. Every decision is finished at compile time; the runtime overhead is zero.

std::invoke_result_t ​

C++17 also gave us a companion, std::invoke_result_t, which obtains the return type of a std::invoke call at compile time. When writing generic code, it has bailed us out more than once:

C++
#include <type_traits>
#include <functional>

template<typename Func, typename... Args>
auto safe_call(Func&& func, Args&&... args)
    -> std::invoke_result_t<Func, Args...>
{
    using Ret = std::invoke_result_t<Func, Args...>;

    if constexpr (std::is_void_v<Ret>) {
        std::invoke(std::forward<Func>(func), std::forward<Args>(args)...);
        std::cout << "(void return)\n";
    } else {
        Ret result = std::invoke(std::forward<Func>(func),
                                 std::forward<Args>(args)...);
        std::cout << "result: " << result << "\n";
        return result;
    }
}

invoke's Performance ​

When we use std::invoke inside template code, the compiler sees the full call chain and inlines it down to the same code as a direct call. We measured it: under -O2, std::invoke calls performed identically to direct calls — a dead heat within measurement error (the occasional slightly-faster reading is just test noise). The reason is plain to see too: std::invoke itself is only a compile-time dispatch wrapper, and after optimization it gets inlined away completely.

Verification code: the performance script we ran is at code/volumn_codes/vol2/ch03-lambda/test_invoke_performance.cpp.

Assembly verification: we generated assembly and had a look (g++ -O2 -S): the direct call, std::invoke, the function pointer, and the lambda all compiled to exactly the same code — the result is computed outright and returned, without a single call instruction left. In the test the call target was known at compile time, which is why the function pointer got folded away along with everything else. If the call target is only decided at runtime, it still goes through an indirect call.

Of course, if you invoke through an object stored in a std::function, the indirection overhead is brought in by std::function's type erasure; it has nothing to do with std::invoke.


Zero-Overhead Callback Design: Templates + Lambdas ​

We already have a firm grip on where std::function's overhead comes from: type erasure, possible heap allocation, indirect calls. So let's think one step further: in many scenarios, the callback's type is nailed down at the moment of registration — do we even need type erasure at all?

No, we don't. The simplest zero-overhead scheme is to pass the lambda directly as a template parameter: the compiler knows the full closure type, and even the call can be fully inlined.

Expand codeCollapse29 lines
C++
#include <algorithm>
#include <vector>
#include <iostream>

// Two template parameters, each taking a callable — zero overhead
template<typename Pred, typename Action>
void for_each_if(std::vector<int>& data, Pred pred, Action action) {
    for (auto& elem : data) {
        if (pred(elem)) {
            action(elem);
        }
    }
}

void demo_template_callback() {
    std::vector<int> data = {1, 2, 3, 4, 5, 6, 7, 8};

    int threshold = 5;
    int sum = 0;

    // Lambdas passed straight to template parameters, fully inlined
    for_each_if(data,
        [threshold](int x) { return x > threshold; },   // predicate
        [&sum](int& x) { sum += x; }                     // action
    );

    std::cout << "Sum of elements > " << threshold << ": " << sum << "\n";
    // Output: Sum of elements > 5: 21
}

But this scheme has its weak spot too: every distinct lambda type instantiates a distinct template function, so you cannot put callbacks of different types into the same container. If the design truly needs runtime polymorphism (say, an event queue holding callbacks of all kinds), some form of type erasure cannot be dodged.

Manual Type Erasure: A Function Pointer Table in Place of Virtual Functions ​

If you need type erasure and also want to dodge std::function's full set of costs, there is one more road: hand-write a lightweight type-erased container. The idea really comes down to two moves — a function pointer table replaces the vtable, and a fixed-size buffer on the stack replaces heap allocation:

Expand codeCollapse77 lines
C++
#include <cstddef>
#include <utility>
#include <iostream>
#include <new>

template<typename Signature, std::size_t BufSize = 32>
class LightCallback;

template<typename R, typename... Args, std::size_t BufSize>
class LightCallback<R(Args...), BufSize> {
    // Operation table: function pointers instead of virtual functions
    struct VTable {
        void (*move)(void* dst, void* src);
        void (*destroy)(void* obj);
        R (*invoke)(void* obj, Args... args);
    };

    // A dedicated VTable generated for each callable type
    template<typename T>
    struct VTableFor {
        static void do_move(void* dst, void* src) {
            new(dst) T(std::move(*static_cast<T*>(src)));
        }
        static void do_destroy(void* obj) {
            static_cast<T*>(obj)->~T();
        }
        static R do_invoke(void* obj, Args... args) {
            return (*static_cast<T*>(obj))(std::forward<Args>(args)...);
        }
        static constexpr VTable value{do_move, do_destroy, do_invoke};
    };

    alignas(std::max_align_t) unsigned char storage_[BufSize];
    const VTable* vtable_ = nullptr;

public:
    LightCallback() = default;

    template<typename T>
    LightCallback(T&& callable) {
        using Decay = std::decay_t<T>;
        static_assert(sizeof(Decay) <= BufSize, "Callable too large for buffer");
        static_assert(alignof(Decay) <= alignof(std::max_align_t),
                     "Callable alignment too high");
        new(storage_) Decay(std::forward<T>(callable));
        vtable_ = &VTableFor<Decay>::value;
    }

    LightCallback(LightCallback&& other) noexcept : vtable_(other.vtable_) {
        if (vtable_) {
            vtable_->move(storage_, other.storage_);
            other.vtable_ = nullptr;
        }
    }

    ~LightCallback() {
        if (vtable_) vtable_->destroy(storage_);
    }

    LightCallback(const LightCallback&) = delete;
    LightCallback& operator=(const LightCallback&) = delete;

    R operator()(Args... args) {
        return vtable_->invoke(storage_, std::forward<Args>(args)...);
    }

    explicit operator bool() const { return vtable_ != nullptr; }
};

void demo_light_callback() {
    int multiplier = 3;
    LightCallback<int(int), 32> cb = [multiplier](int x) {
        return x * multiplier;
    };

    std::cout << cb(14) << "\n";  // 42
}

Put this LightCallback next to std::function and, yes, it falls a notch short on generality — copy support and allocator support are simply not there. But it catches the most common use cases: storing lambdas with captures, skipping the heap allocation, and leaving only one layer of indirection.

Selection Guide ​

Back to the event system from the opening, and we can lay the options out. If your callback needs no context and sits on a hot path, the function pointer is the least trouble; the targets it can point at are captureless lambdas, plain functions, and the like. If the event queue wants runtime polymorphism, std::function steps up — and the cost is equally plain: even when the object lands inside the SBO range, the layer of indirection that comes with the vtable blocks inlining, and in our measurements it ran 7-9x slower than a direct call. If the callback's type is settled at compile time, go straight to template parameters — what you get is zero overhead, only they cannot go into containers. If you want runtime polymorphism and performance both, hand-written type erasure is the road; it takes a bit more code, and the behavior becomes entirely yours to control.

Performance data source: still our same setup, code/volumn_codes/vol2/ch03-lambda/test_function_performance.cpp, run on GCC 15.2.1 with -O2, 100 million calls in total.


References ​

pdf-latest-4-g85128cc · 85128cc · 2026-10-05