Concurrency & Parallelism
std::async. Compile every example with -pthread on GCC/Clang.
Threads
A thread is an independent flow of execution inside one process. All threads in a process
share the same memory, which is exactly what makes them fast — and exactly what makes them
dangerous. You create a thread with std::thread, passing it a callable and its
arguments, and you must join() it (wait for it) before it is destroyed.
#include <thread>
#include <iostream>
void work(int id) {
std::cout << "worker " << id << " running\n";
}
int main() {
std::thread t1(work, 1); // starts running immediately
std::thread t2(work, 2);
t1.join(); // block until t1 finishes
t2.join(); // block until t2 finishes
}
If you destroy a std::thread that is still joinable (never joined and never
detached), the program calls std::terminate. Forgetting to join is the single most
common threading bug — the RAII habit below fixes it.
Creating and joining
Prefer std::jthread when your compiler is C++20: it joins automatically in its
destructor and supports cooperative cancellation. Until then, wrap std::thread in a
small RAII guard so a thrown exception can never leave a thread un-joined.
class ThreadGuard { // joins on scope exit, like a smart pointer
public:
explicit ThreadGuard(std::thread &t) : t_(t) {}
~ThreadGuard() { if (t_.joinable()) t_.join(); }
ThreadGuard(const ThreadGuard &) = delete;
ThreadGuard &operator=(const ThreadGuard &) = delete;
private:
std::thread &t_;
};
void safe() {
std::thread t(work, 7);
ThreadGuard guard(t); // if anything throws, the thread is still joined
std::cout << "doing more work\n";
} // guard's destructor joins here
Sharing data safely
When two threads touch the same variable without synchronisation you have a data race, and a data race is undefined behaviour — the result is not merely "wrong", the compiler may legally do anything. The example below increments a shared counter from four threads; without a lock the total comes out too small.
long total = 0; // SHARED — must be synchronised
void addUnsafe(int rounds) {
for (int i = 0; i < rounds; ++i)
++total; // read-modify-write: three steps, not one
}
// Four threads each doing 100'000 rounds give a total well below 400'000,
// because updates overwrite each other.
Synchronization
To make a shared update safe you must make it atomic — one thread completes the whole read-modify-write before another begins. C++ gives you two families of tools: locks (mutexes) that guard a region of code, and atomics that make a single variable indivisible.
Mutexes and lock guards
A std::mutex is a lock. Wrap every access to the shared data in a
std::lock_guard — an RAII object that unlocks in its destructor, so you cannot forget
to release the lock even if an exception is thrown.
#include <mutex>
std::mutex m;
long total = 0;
void addSafe(int rounds) {
for (int i = 0; i < rounds; ++i) {
std::lock_guard<std::mutex> lock(m); // locks here...
++total; // ...read-modify-write is now safe
} // ...and unlocks at the closing brace
}
For a critical section that must lock more than one mutex, use std::scoped_lock
(C++17) — it locks all of them without deadlock. std::unique_lock is the flexible
variant you need when you must unlock and re-lock manually, or hand the lock to a condition variable.
Atomics
When only a single variable needs protection, an std::atomic is usually faster and
simpler than a mutex: its operations are indivisible and lock-free where the hardware supports it.
#include <atomic>
std::atomic<long> count{0};
void bump(int rounds) {
for (int i = 0; i < rounds; ++i)
count.fetch_add(1, std::memory_order_relaxed); // one indivisible step
}
The default memory order (seq_cst) is the safest and easiest to reason about;
relaxed is enough for a plain counter. Reach for the weaker orders only when you have
measured a need and understand the visibility rules — they trade safety for speed.
| Tool | Protects | Use when |
|---|---|---|
std::mutex + lock_guard | a region of code / several variables | invariants spanning more than one value |
std::scoped_lock | many mutexes at once | you must hold two locks together, deadlock-free |
std::atomic | one variable | counters, flags, single-word hand-off |
std::condition_variable | a wait/notify protocol | a thread waits for a state change |
Tasks & futures
Raw threads force you to manage lifetimes by hand. The task-based model is higher level: you ask
the standard library to run a function and hand you a std::future — a placeholder for
a result that will be ready later. std::async is the simplest way to get one.
#include <future>
#include <numeric>
double heavy(int n) { /* ... expensive work ... */ return 0.0; }
int main() {
std::future<double> job = std::async(std::launch::async, heavy, 1'000'000);
// ... do unrelated work on the main thread while `job` runs ...
double result = job.get(); // blocks until the task is finished, returns the value
}
get() may be called only once. Behind the scenes a future is fed by a
promise or a packaged_task; those are the building blocks of thread pools and
work queues, where many tasks are executed by a fixed set of worker threads.
Parallel patterns
The most common real pattern is "map": apply a function to every element of a range, in parallel. Split the range into chunks, hand each chunk to a worker, then join. Keeping each chunk to its own slice of the output means the workers never collide and no locking is needed.
#include <thread>
#include <vector>
// Out-of-place parallel map: out[i] = f(in[i]); workers write disjoint ranges.
void parallelMap(const std::vector<int> &in, std::vector<int> &out,
int (*f)(int), int workers) {
out.resize(in.size());
std::vector<std::thread> pool;
const std::size_t n = in.size();
for (int w = 0; w < workers; ++w) {
std::size_t begin = n * w / workers; // this worker's slice
std::size_t end = n * (w + 1) / workers;
pool.emplace_back([&, begin, end] {
for (std::size_t i = begin; i < end; ++i) out[i] = f(in[i]);
});
}
for (auto &t : pool) t.join();
}
Because each thread writes a different region of out, there is no shared mutation and
therefore no data race. C++17 goes further with execution policies: a single line requests
a parallel standard algorithm without you writing any threading at all.
#include <algorithm>
#include <execution>
std::for_each(std::execution::par, data.begin(), data.end(), [](int &x) {
x = x * x; // may run on multiple threads
});
Pitfalls
Concurrency bugs are subtle because they hide in timing. Keep this list in mind and most of them will never reach production.
- Data races are undefined behaviour. "It worked on my machine" is meaningless; use a lock or an atomic.
- Deadlock. Two threads each waiting for the other's lock. Lock mutexes in one global order, or use
std::scoped_lock. - Un-joined threads abort the program when a joinable
std::threadis destroyed — alwaysjoin()(or use a guard /std::jthread). - Over-locking destroys scalability. Shrink critical sections; prefer private data and merge at the end.
- False sharing. Two threads writing different variables that sit on the same cache line still fight; separate them.
- Threads are expensive. Do not spawn one per item — use a fixed pool sized to the core count.
Practice
- Run the data-race example in demo 11, then add the mutex and watch the total become exact.
- Convert the mutex counter to
std::atomic<long>and compare the performance. - Parallelise a sum over a large vector by giving each thread its own partial sum, then adding the partials at the end.
- Add a
ThreadGuardto a function that throws, and prove the thread is still joined. - Start three
std::asynctasks and collect their results withfuture::get().