Debugging Techniques for Concurrent Programs
Let's be honest: the pain of debugging concurrent programs can only truly be understood by someone who has personally stepped in that minefield. Bugs in single-threaded programs are at least deterministic—give them the same input, and they will blow up in the same place, in the same way, every time. Concurrency bugs are not like that. A data race might show up once in ten thousand runs; a deadlock might only trigger under one particular thread scheduling order; and it is always "works on my machine, fails on CI without fail". We once spent two full days on a data race that turned out to be a lambda capturing a reference to a local variable—the kind of bug you simply cannot see by reading the code, because on a single-threaded execution path it is perfectly correct.
What we will build in this article is a systematic methodology for debugging concurrency. Not the "add a print and see what happens" kind, but a full pipeline: start from understanding the characteristic signatures of each bug class, move on to selecting the right tool, then learn to read the tool's report, and end up with a fix process that is reproducible and verifiable. We will focus on ThreadSanitizer (TSan), Valgrind's Helgrind tool, Clang's compile-time thread safety analysis, and a practical structured logging scheme.
Environment Notes
All commands and code in this article have been tested on the following environment: Ubuntu 22.04 LTS (WSL2 works too), with Clang 16+ or GCC 12+ as the compiler (TSan support required), Valgrind 3.18 or newer (a plain apt install valgrind does it), GDB 12+ as the debugger, and CMake 3.20+ if that is what manages your project. On older distributions the exact TSan report format may differ slightly, but the core content is the same.
The Four Schools of Concurrency Bugs
Before reaching for tools, we first need a rough map of how concurrency bugs divide into categories, because different bug types call for completely different diagnostic strategies.
Data races are the most common and the most insidious kind. The definition is strict: two or more threads access the same memory location concurrently, at least one of the accesses is a write, and there is no synchronization relationship between them (no mutex, no atomic, no happens-before). The C++ standard explicitly declares a data race to be undefined behavior—not "might go wrong", but "anything can happen", including but not limited to reading garbage values, crashing, or even appearing to "work correctly" and then suddenly detonating one day. Data races are hard to track because they depend on the thread scheduling order, and that order under your debugger may be completely different from the one in production. You add a printf to debug something, the printing itself changes the timing, and the bug vanishes—this is the classic "Heisenbug".
Deadlock is the other big category. Two or more threads each wait for resources held by the other; nobody yields, and the program locks up completely. Deadlocks are actually more deterministic than data races—once the specific lock acquisition order is triggered, it is guaranteed to happen. The problem is that the trigger conditions can be extremely complex, involving particular combinations of execution paths across multiple threads. Worse, deadlocks often refuse to appear under normal load and only surface under certain concurrency patterns.
Livelock is sneakier than deadlock. The threads are not stuck—CPU usage may be at 100%—yet no meaningful progress happens. A classic example: two threads politely keep yielding the resource to each other, and neither ever gets it. A livelock manifests as the program slowing down rather than freezing, so it is easy to misdiagnose as a performance problem.
Finally, there is the dangling reference. A thread accesses an object through a reference or pointer that has already outlived its lifetime—this is especially common in asynchronous programming. Say you start a thread, pass in a reference to a local variable, the function returns, the local variable is destroyed, and the thread is still using that reference. How this bug manifests depends on what that memory gets reallocated for: it may read a value that "looks normal but is actually wrong", or it may segfault outright.
| Bug type | Core signature | Reproduction difficulty | Typical signals |
|---|---|---|---|
| Data race | Unsynchronized concurrent read/write | Very high (timing-dependent) | Intermittent wrong results, Heisenbugs |
| Deadlock | Circular wait for resources | Medium-high (path-dependent) | Program frozen solid |
| Livelock | Repeated yielding with no progress | Medium | CPU at 100% but no output |
| Dangling reference | Accessing a destroyed object | High (memory-state-dependent) | Intermittent crashes, garbage values |
ThreadSanitizer: The Nemesis of Data Races
How It Works: Compiler Instrumentation
ThreadSanitizer (TSan for short) works by instrumenting your code at compile time. When you add the -fsanitize=thread compile option, the compiler inserts extra checking code around every memory access (both reads and writes). At runtime, this checking code maintains a "shadow memory" that records the access history and synchronization events of every memory location.
TSan uses a hybrid algorithm based on happens-before relationships plus lockset analysis. In short, it tracks the thread ID and a logical timestamp (a vector clock) for every memory access, while also tracking which mutexes the current thread holds. If it sees two memory accesses from different threads with no happens-before relationship between them (that is, no synchronization operation of any kind), and at least one of them is a write, it reports a data race. The theory behind the algorithm guarantees that if a data race actually occurs during your test run (even exactly once), the algorithm is guaranteed to detect it. One caveat, though: TSan's implementation maintains a bounded history buffer per 8-byte memory location, so in extreme cases (say, a large number of threads frequently hitting the same address until old records get evicted) the real miss rate is very low but not zero. For the overwhelming majority of real-world scenarios, you can treat "TSan reported nothing" as a strong signal that "there really is no data race on this execution path".
Enabling TSan
Enabling TSan is dead simple—just add the corresponding flag at compile time:
# Add -fsanitize=thread and debug info when compiling
clang++ -fsanitize=thread -g -O1 -pthread your_program.cpp -o your_program
# Or with GCC
g++ -fsanitize=thread -g -O1 -pthread your_program.cpp -o your_programA few things to watch out for here. First, -g is mandatory; without it, the TSan report contains only addresses and no source locations, which makes pinpointing the problem painful. Second, the official recommendation is -O1 or higher, mostly for performance—TSan already imposes a 5-15x slowdown, and the unoptimized -O0 code piles snow onto that avalanche; but do not go to -O2 or above either, because aggressive optimization can inline so many functions that the stack traces become hard to read. Third, TSan cannot be used together with AddressSanitizer (ASan); if your build script turns both on at once, the compiler will error out outright.
If you use CMake, you can configure it like this:
# Enable TSan in CMakeLists.txt
option(ENABLE_TSAN "Enable ThreadSanitizer" OFF)
if(ENABLE_TSAN)
add_compile_options(-fsanitize=thread -g -O1)
add_link_options(-fsanitize=thread)
endif()Then a plain cmake -DENABLE_TSAN=ON .. does it.
Hands-On: A Complete Data Race Diagnosis
Let's look at a classic data race scenario, then hunt it down with TSan step by step.
Expand codeCollapse42 lines
#include <thread>
#include <vector>
#include <iostream>
class ThreadSafeCounter {
public:
void increment()
{
// Looks completely harmless, but there is a data race here
count_++;
}
int get() const { return count_; }
private:
int count_{0};
};
int main()
{
ThreadSafeCounter counter;
constexpr int kIterations = 100000;
auto worker = [&counter]() {
for (int i = 0; i < kIterations; ++i) {
counter.increment();
}
};
std::vector<std::thread> threads;
for (int i = 0; i < 4; ++i) {
threads.emplace_back(worker);
}
for (auto& t : threads) {
t.join();
}
// We expect 400000; in practice you will almost never get it
std::cout << "Final count: " << counter.get() << "\n";
return 0;
}The problem in this code is obvious—count_++ is not an atomic operation, and four threads incrementing it concurrently will lose updates. The catch is that without TSan, all you see is a "wrong result" (say, 287541 instead of 400000), and you cannot tell whether it is a data race or a logic error. Add TSan:
clang++ -fsanitize=thread -g -O1 -pthread counter.cpp -o counter
./counterTSan's output looks roughly like this (exact line numbers will vary with your code):
Expand codeCollapse22 lines
==================
WARNING: ThreadSanitizer: data race (pid=12345)
Write of size 4 at 0x7b0c00000000 by thread T2:
#0 ThreadSafeCounter::increment() counter.cpp:10:9 (counter+0x4a2b)
#1 main::$_0::operator()() const counter.cpp:24:13 (counter+0x4a03)
Previous write of size 4 at 0x7b0c00000000 by thread T1:
#0 ThreadSafeCounter::increment() counter.cpp:10:9 (counter+0x4a2b)
#1 main::$_0::operator()() const counter.cpp:24:13 (counter+0x4a03)
Location is stack of main thread.
Thread T2 (tid=12347, running) created by main thread at:
#0 pthread_create <null> (counter+0x42278)
#1 main counter.cpp:28:23 (counter+0x4b0e)
Thread T1 (tid=12346, finished) created by main thread at:
#0 pthread_create <null> (counter+0x42278)
#1 main counter.cpp:28:23 (counter+0x4b0e)
SUMMARY: ThreadSanitizer: data race counter.cpp:10:9 in ThreadSafeCounter::increment()
==================
Final count: 287541Let's dissect this report. The top line, WARNING: ThreadSanitizer: data race, tells you this is a data race. Then it presents the two conflicting accesses: one is a write by thread T2 (Write of size 4), occurring at counter.cpp:10:9—the count_++ line. The other is the previous write by thread T1 (Previous write), at the same location. This is precisely the textbook definition of a data race—two threads writing the same memory location with no synchronization. Finally, it tells you where the threads were created (main counter.cpp:28:23), helping you trace the whole call chain.
The fix is simple—use std::atomic<int>, or add a mutex:
#include <atomic>
class ThreadSafeCounter {
public:
void increment()
{
// Use an atomic; the data race is gone
count_.fetch_add(1, std::memory_order_relaxed);
}
int get() const
{
return count_.load(std::memory_order_relaxed);
}
private:
std::atomic<int> count_{0};
};Recompile and run: TSan no longer reports anything, and the output settles at a stable 400000.
TSan's Limitations
TSan is great to use, but it is not omnipotent, and we must be clear about its limits.
First, the performance cost is heavy. TSan's typical overhead is a 5-15x runtime slowdown plus 5-10x memory overhead. That means you cannot run with TSan enabled in production—it is for testing and CI only. The good news is that you do not need it in production, because TSan detects problems in code logic, not problems in the runtime environment.
Second, TSan can only detect data races on the code paths that your tests actually execute. If your test coverage is insufficient, some races may never fire. So when using TSan, your concurrency tests should cover as many thread interleavings as possible—for example, run several rounds with different thread counts and different task granularities.
There is also an easily overlooked issue: TSan's recognition of custom synchronization mechanisms is limited. If you implement your own spinlock or barrier based on std::atomic but do not use the annotation interface TSan provides (__tsan_acquire / __tsan_release), TSan may produce false positives (treating your custom synchronization as no synchronization) or miss races. For the standard std::mutex, std::atomic, std::condition_variable, and friends, TSan recognizes them all correctly; but if you have custom synchronization primitives, extra handling is required.
⚠️ Note: TSan and ASan cannot be enabled at the same time. If your project already uses ASan for memory error detection, you need to build a separate TSan version. The usual practice is to run two test suites in CI—one with ASan, one with TSan.
Helgrind: Valgrind's Thread Error Detector
How It Works and How to Use It
Helgrind is a thread error detector in the Valgrind toolset. Unlike TSan's compile-time instrumentation, Valgrind uses dynamic binary instrumentation (DBI)—it does not require recompiling your program; instead it dynamically analyzes every single instruction at runtime.
Helgrind uses happens-before based lockset analysis. It tracks every pthread synchronization operation in the program (mutex lock/unlock, thread create/join, condition variable signal/wait) and builds a happens-before relationship graph between threads. At the same time, it maintains a "lockset" for each thread (the set of locks currently held) and checks on every memory access: if two threads access the same memory location and their locksets have an empty intersection (meaning no common lock protects it), it reports a potential data race.
Additionally, Helgrind builds a "lock order graph". If it observes lock A acquired before lock B (forming an A -> B edge), and later observes a B -> A order in another thread, a cycle appears in the graph—and that is a potential deadlock.
Using Helgrind requires no recompilation—just run it directly:
valgrind --tool=helgrind ./your_programIf your program takes command-line arguments, just append them at the end:
valgrind --tool=helgrind ./your_program --arg1 --arg2Hands-On: A Lock Order Error
Here is a classic lock order problem—two threads acquiring two locks in different orders. This is prime breeding ground for deadlocks.
Expand codeCollapse57 lines
#include <mutex>
#include <thread>
#include <iostream>
class BankAccount {
public:
explicit BankAccount(int balance) : balance_(balance) {}
void transfer_from(BankAccount& other, int amount)
{
// Lock self first, then the counterparty
std::lock_guard<std::mutex> lk1(mutex_);
std::lock_guard<std::mutex> lk2(other.mutex_);
if (other.balance_ >= amount) {
other.balance_ -= amount;
balance_ += amount;
}
}
int get_balance() const
{
std::lock_guard<std::mutex> lk(mutex_);
return balance_;
}
private:
mutable std::mutex mutex_;
int balance_;
};
int main()
{
BankAccount alice(1000);
BankAccount bob(1000);
// alice transfers 100 to bob
std::thread t1([&]() {
for (int i = 0; i < 100; ++i) {
alice.transfer_from(bob, 1);
}
});
// bob transfers 100 to alice
std::thread t2([&]() {
for (int i = 0; i < 100; ++i) {
bob.transfer_from(alice, 1);
}
});
t1.join();
t2.join();
std::cout << "Alice: " << alice.get_balance()
<< ", Bob: " << bob.get_balance() << "\n";
return 0;
}This program can deadlock with some probability: t1 locks alice then bob, while t2 locks bob then alice. If t1 locks alice's mutex at the same moment t2 locks bob's, both end up waiting for the other to release—the classic deadlock. Run it under Helgrind:
g++ -g -O1 -pthread transfer.cpp -o transfer
valgrind --tool=helgrind ./transferHelgrind will output a report like this:
Expand codeCollapse28 lines
---Thread-Announcement ---
Thread #1 is the program's root thread
---Thread-Announcement ---
Thread #2 was created
at 0x4C3A0E3: pthread_create (helgrind_intercepts.c:xxx)
by 0x401234: main (transfer.cpp:38)
---Thread-Announcement ---
Thread #3 was created
...
--- Lock order violation ---
Possible data race during lock order check
Lock #1 (0x....) locked at
...
by 0x4011A0: BankAccount::transfer_from (transfer.cpp:13)
Lock #2 (0x....) locked at
...
by 0x4011C8: BankAccount::transfer_from (transfer.cpp:14)
Lock #2 (0x....) previously locked at
...
by 0x401208: main::$_1::operator() (transfer.cpp:44)
Lock #1 (0x....) previously locked at
...
by 0x401208: main_$_1::operator() (transfer.cpp:44)
This indicates that the locking order is inconsistent.Helgrind tells you plainly: the lock acquisition order is inconsistent. One path takes #1 before #2 (transfer.cpp:13-14); the other path takes #2 before #1. The fix is to use std::lock to acquire both locks together; internally it uses a try-and-back-off algorithm that avoids deadlock:
void transfer_from(BankAccount& other, int amount)
{
// std::lock acquires both locks together, avoiding deadlock
std::lock(mutex_, other.mutex_);
std::lock_guard<std::mutex> lk1(mutex_, std::adopt_lock);
std::lock_guard<std::mutex> lk2(other.mutex_, std::adopt_lock);
if (other.balance_ >= amount) {
other.balance_ -= amount;
balance_ += amount;
}
}TSan vs Helgrind: Which to Choose
These two tools overlap in quite a few capabilities, but each has its own focus.
TSan is compile-time instrumentation: it requires recompilation but has relatively modest runtime overhead (though still a 5-15x slowdown), it has the best support for C++ standard library synchronization primitives, and its report format is clean and readable. If you can recompile the project, TSan is usually the first choice—its data race detection is more precise, with a lower false positive rate.
Helgrind is runtime dynamic analysis: no recompilation needed (as long as debug symbols are present), but its runtime overhead is even larger than TSan's (typically a 20-50x slowdown), because every instruction has to be translated through Valgrind's IR. Helgrind's advantage is that you can take an already-compiled binary and analyze it directly, without setting up a build environment. Also, Helgrind's lock order analysis is particularly strong—if you suspect a deadlock risk that has not triggered yet, Helgrind's lock order graph can surface the hazard ahead of time.
Our advice: use TSan in daily development for quickly catching data races; bring in Helgrind when you need to analyze lock order problems or cannot recompile. The two complement each other—there is no need to pick just one.
A Compile-Time Line of Defense: Clang Thread Safety Analysis
TSan and Helgrind are both runtime tools—the bug has to actually happen before they can detect it. But one class of problems can be stopped at compile time. Clang's Thread Safety Analysis (TSA) is a compile-time static analysis extension: you declare thread safety constraints through code annotations, and the compiler then checks at build time whether you have violated those constraints. Zero runtime overhead, zero performance impact—it works entirely at compile time.
Basic Annotations
TSA's core concept is the "capability". A mutex is a capability—holding it is what entitles you to access the data it protects. You declare these constraints with macros (backed by __attribute__ underneath).
First, add the CAPABILITY annotation to your mutex type:
Expand codeCollapse23 lines
// Wrap the standard library mutex in an annotated type
class CAPABILITY("mutex") Mutex {
public:
void lock() ACQUIRE() { mu_.lock(); }
void unlock() RELEASE() { mu_.unlock(); }
bool try_lock() TRY_ACQUIRE(true) { return mu_.try_lock(); }
private:
std::mutex mu_;
};
// The RAII guard needs annotations too
class SCOPED_CAPABILITY MutexGuard {
public:
explicit MutexGuard(Mutex& m) ACQUIRE(m) : mu_(m) { mu_.lock(); }
~MutexGuard() RELEASE() { mu_.unlock(); }
MutexGuard(const MutexGuard&) = delete;
MutexGuard& operator=(const MutexGuard&) = delete;
private:
Mutex& mu_;
};Then you can use GUARDED_BY to declare which mutex protects a data member, and REQUIRES to declare which lock must be acquired before a function is called:
Expand codeCollapse32 lines
class ThreadSafeQueue {
public:
void push(int value)
{
MutexGuard lk(mutex_); // Acquire the lock
data_.push_back(value); // OK, lock held
}
int pop()
{
MutexGuard lk(mutex_);
int val = data_.front(); // OK
data_.pop_front();
return val;
}
// Dangerous! Reads directly, skipping the lock
int unsafe_front()
{
return data_.front(); // Compiler warning!
}
// Declare that the caller must hold the lock
int front_locked() REQUIRES(mutex_)
{
return data_.front(); // OK, the caller guarantees the lock is held
}
private:
mutable Mutex mutex_;
std::deque<int> data_ GUARDED_BY(mutex_);
};Compile with -Wthread-safety, and unsafe_front() triggers a compiler warning, because it accesses data_—protected by GUARDED_BY(mutex_)—without holding mutex_. front_locked(), on the other hand, carries the REQUIRES(mutex_) annotation, so the compiler knows it requires the caller to hold the lock and that accessing data_ inside is safe—if someone calls front_locked() without the lock, the warning shows up on the caller's side.
Lock Order Annotations
TSA also supports declaring lock acquisition order, to prevent deadlocks:
class NetworkManager {
private:
Mutex stats_mutex_ ACQUIRED_AFTER(data_mutex_);
Mutex data_mutex_;
std::vector<int> data_ GUARDED_BY(data_mutex_);
int total_bytes_ GUARDED_BY(stats_mutex_);
};If somewhere you lock data_mutex_ first and then stats_mutex_, no problem—that matches the declared order. But reverse it—locking stats_mutex_ first, then data_mutex_—and the compiler sounds the alarm.
Enabling it is simple:
clang++ -Wthread-safety -c your_file.cpp⚠️ Note: TSA is pure static analysis and cannot replace runtime tools. It only checks the constraints you annotated; code without annotations gets zero scrutiny from it. And TSA is currently a Clang-exclusive extension—GCC and MSVC do not support it. But if you build with Clang, annotating key data structures and letting the compiler stand guard for you saves an enormous amount of debugging time.
Runtime Diagnosis of Deadlocks
TSA can prevent some deadlocks at compile time, but if your program is already frozen, you need runtime diagnostic methods.
GDB: The Most Direct Approach
When the program deadlocks, the most direct move is to attach GDB to the process and inspect the call stacks of all threads:
# Find your program's PID
ps aux | grep your_program
# Attach GDB
gdb -p <PID>
# In GDB: show the call stacks of all threads
(gdb) thread apply all btYou will see output like this:
Thread 3 (Thread 0x7f... "your_program"):
#0 __lll_lock_wait (futex=..., private=0) at lowlevellock.c:52
#1 __pthread_mutex_lock (mutex=...) at pthread_mutex_lock.c:67
#2 BankAccount::transfer_from (this=..., other=..., amount=1) at transfer.cpp:13
#3 ...
Thread 2 (Thread 0x7f... "your_program"):
#0 __lll_lock_wait (futex=..., private=0) at lowlevellock.c:52
#1 __pthread_mutex_lock (mutex=...) at pthread_mutex_lock.c:67
#2 BankAccount::transfer_from (this=..., other=..., amount=1) at transfer.cpp:13
#3 ...Both threads are stuck in __lll_lock_wait (the kernel-side wait of the mutex), and both at line 13 of transfer_from—ironclad evidence of a deadlock. From the call stacks you can infer the lock acquisition order, and then fix it.
A GDB Python Script to Help
For larger projects, reading thread apply all bt output by eye is exhausting. You can write a small GDB Python script that extracts every thread waiting on a lock, along with the mutex address it waits for:
Expand codeCollapse25 lines
# save as deadlock_detector.py
import gdb
class DeadlockDetector(gdb.Command):
"""Detect potential deadlocks by showing all threads waiting on mutexes."""
def __init__(self):
super().__init__("detect-deadlock", gdb.COMMAND_USER)
def invoke(self, arg, from_tty):
threads = gdb.selected_inferior().threads()
for thread in threads:
thread.switch()
frame = gdb.selected_frame()
sal = frame.find_sal()
try:
# Try to find __lll_lock_wait or pthread_mutex_lock
func_name = frame.function().name or ""
if "lock" in func_name.lower():
print(f"Thread {thread.num} waiting on lock at "
f"{sal.symtab.filename}:{sal.line}")
except Exception:
pass
DeadlockDetector()After source deadlock_detector.py inside GDB, just type detect-deadlock and you will see every thread waiting on a lock.
Structured Logging: Making printf a Bit More Reliable
When debugging concurrent programs, many people's first instinct is to add printf or std::cout calls. That has two serious problems.
First, printf and std::cout are not themselves thread-safe (strictly speaking, the C++ standard guarantees they do not cause data races, but when multiple threads write to std::cout simultaneously, the output interleaves into chaos). You add a pile of prints, and what you see may be one line garbled by another thread's output cutting in mid-line—worse than having no logs at all.
Second, logs without timestamps and thread identifiers are nearly useless. When you see two lines of output, value = 42 and value = 0, you have no idea which thread wrote which one, when, or in what order.
A Minimal Thread-Safe Logger
What we need is a thread-safe logger where every entry carries a timestamp and a thread ID. The implementation below is simple but practical:
Expand codeCollapse43 lines
#include <mutex>
#include <chrono>
#include <sstream>
#include <iostream>
#include <thread>
#include <iomanip>
#include <atomic>
class ThreadSafeLogger {
public:
static ThreadSafeLogger& instance()
{
static ThreadSafeLogger logger;
return logger;
}
void log(const std::string& level, const std::string& message)
{
auto now = std::chrono::steady_clock::now();
auto ns = std::chrono::duration_cast<std::chrono::nanoseconds>(
now.time_since_epoch());
// Build the complete log line locally first, then lock and output in one shot
// This keeps the lock hold time to a minimum
std::ostringstream oss;
oss << "[" << std::setw(16) << ns.count() << " ns] "
<< "[" << std::this_thread::get_id() << "] "
<< "[" << level << "] "
<< message << "\n";
std::lock_guard<std::mutex> lk(mutex_);
std::cout << oss.str();
}
private:
ThreadSafeLogger() = default;
std::mutex mutex_;
};
// Convenience macros to save typing
#define LOG_INFO(msg) ThreadSafeLogger::instance().log("INFO", msg)
#define LOG_WARN(msg) ThreadSafeLogger::instance().log("WARN", msg)
#define LOG_ERROR(msg) ThreadSafeLogger::instance().log("ERROR", msg)The key implementation detail: we first build the complete log line in a local std::ostringstream, and only then take the lock and output. The benefit is that the lock is held for a very short time (just one std::cout << string), which reduces lock contention. If you do the formatting inside the lock, multiple threads end up queuing up behind the formatting, and the impact on concurrent performance is not negligible.
Every log entry carries three pieces of key information: a nanosecond-resolution timestamp (for establishing event order), the thread ID (for telling different threads' behavior apart), and the log level. With these, when analyzing a concurrency bug you can trace each thread's timeline precisely.
Using it is simple:
ThreadSafeLogger::instance().log("INFO",
"Acquired mutex for account " + std::to_string(account_id));The output looks like this:
[ 123456789012345 ns] [140234567890] [INFO] Acquired mutex for account 42
[ 123456789045678 ns] [140234567891] [INFO] Acquired mutex for account 17From the timestamps and thread IDs you can clearly see two threads acquiring different mutexes almost simultaneously—if they later go after a second mutex in opposite orders, you have found the root cause of a deadlock.
⚠️ Note: This logger uses
std::coutas its underlying output. If your program needs high-throughput logging (say, millions of entries per second), this implementation will not cut it—you would need a lock-free ring buffer design, or an off-the-shelf logging library (spdlog and friends). For the debugging phase, though, it is entirely adequate.
A Systematic Diagnostic Workflow
Alright—at this point we have covered the four main tools: TSan, Helgrind, Clang TSA, and structured logging. The question is, when you hit a concurrency bug in a real project, in what order should you use these tools? Based on our own history of stepping in pits, here is the workflow we have settled on.
When you spot a suspected concurrency bug, step one is always to reproduce it as reliably as possible. This is the hardest and the most critical step. Record every condition that triggers the bug: input data, thread count, system load, even the hardware model. If the bug only appears under high concurrency, write a stress test and run it repeatedly; if it only appears with particular data, keep that data. A bug that cannot be reproduced reliably is nearly impossible to fix—because you cannot verify whether your fix works. If it truly cannot be reproduced reliably, consider a loop test in CI—run the same test 1000 times, and a single failure counts as a failure.
Once it reproduces, step two is to classify the bug. Data race, deadlock, livelock, or dangling reference? If the program outputs wrong results without crashing, most likely a data race. If the program freezes solid, probably a deadlock. If the CPU sits at 100% with no output, possibly a livelock. If it segfaults and the stack trace shows strange addresses, suspect a dangling reference. This classification decides which tool you reach for next.
Step three: pick and run the tool. For a data race, build a TSan version and run it. For deadlock risk, use Helgrind's lock order analysis. For a process that is already deadlocked, attach GDB and inspect all thread stacks. For a dangling reference, ASan fits better (this article is mostly about concurrency tools, but ASan's use-after-free detection is extremely precise).
Step four: analyze the tool's report. TSan's report tells you exactly which line of code is problematic and which threads are in conflict. Helgrind tells you where the lock acquisition order is inconsistent. GDB tells you where each thread is stuck. Read the report carefully—do not rush to change the code; first make sure you understand the root cause of the problem.
Step five: fix and verify. After the fix, rerun TSan/Helgrind to confirm the reports are gone, and rerun your reproduction test to confirm the bug no longer appears. If conditions allow, add a TSan build to CI as a standing check, to keep the same class of problem from being reintroduced.
The workflow looks simple, but every step has pitfalls. The most common mistake is skipping "reproduce" and jumping straight into reading code to guess the bug's location—in a concurrent program, your guess is more likely wrong than right, because the root cause of a concurrency bug often hides on a seemingly unrelated code path. Another common mistake is not rerunning TSan to verify after the fix—you believe it is fixed, but in reality you may have merely changed the timing so the bug appears less often, instead of eliminating it at the root.
Where We Are
In this article we assembled a toolbox and a methodology for debugging concurrency. TSan catches data races at runtime through compile-time instrumentation; Helgrind detects lock order problems and races through dynamic analysis; Clang TSA prevents thread safety violations at compile time with annotations; GDB provides a live snapshot when the program is deadlocked; and structured logging helps us trace event timelines during debugging. Each tool has its own focus, and used in combination they cover the vast majority of concurrency bug scenarios.
But "correct" is only half of concurrent programming. A bug-free concurrent program is not automatically an efficient one—you may spend a week optimizing a mutex only to find the bottleneck was never there; or you may chase lock-free performance and end up with code too complex to maintain. The next article discusses how to measure the performance of concurrent programs scientifically: multithreaded usage of Google Benchmark, common pitfalls in concurrent benchmark design, and performance counter analysis with the perf tool. Debugging tells us "what is wrong"; benchmarking tells us "what is slow"—combined, they make up the complete engineering capability for concurrency.
💡 The complete example code lives in Tutorial_AwesomeModernCPP; browse to
code/volumn_codes/vol5/ch08-debug-testing-perf/.
References
- ThreadSanitizer — LLVM Documentation — The official TSan documentation, covering usage, limitations, and configuration options
- Dynamic Race Detection with LLVM Compiler — Google Research — The original TSan-LLVM paper, describing the hybrid detection algorithm in detail
- Helgrind: an experimental thread error detector — Valgrind Manual — The official Helgrind manual, including lock order analysis and the annotation API
- Thread Safety Analysis — Clang Documentation — The complete Clang TSA reference, covering the usage of every annotation
- Thread Safety Analysis in C and C++ — CERT/SEI (CMU) — The design philosophy behind TSA and its industrial applications
- C/C++ Thread Safety Analysis — Google Research (PDF) — The original TSA paper, by Hutchins et al.