Concurrency & Parallelism

Scope: Modern machines have many cores, and C++ lets you use them with the standard library — no extensions required. This lesson runs from creating a thread, through the data races threads cause, to the tools that fix them: mutexes, atomics, and task-based std::async. Compile every example with -pthread on GCC/Clang.

Threads

A thread is an independent flow of execution inside one process. All threads in a process share the same memory, which is exactly what makes them fast — and exactly what makes them dangerous. You create a thread with std::thread, passing it a callable and its arguments, and you must join() it (wait for it) before it is destroyed.

#include <thread>
#include <iostream>

void work(int id) {
    std::cout << "worker " << id << " running\n";
}

int main() {
    std::thread t1(work, 1);   // starts running immediately
    std::thread t2(work, 2);
    t1.join();                 // block until t1 finishes
    t2.join();                 // block until t2 finishes
}

If you destroy a std::thread that is still joinable (never joined and never detached), the program calls std::terminate. Forgetting to join is the single most common threading bug — the RAII habit below fixes it.

Creating and joining

Prefer std::jthread when your compiler is C++20: it joins automatically in its destructor and supports cooperative cancellation. Until then, wrap std::thread in a small RAII guard so a thrown exception can never leave a thread un-joined.

class ThreadGuard {                 // joins on scope exit, like a smart pointer
public:
    explicit ThreadGuard(std::thread &t) : t_(t) {}
    ~ThreadGuard() { if (t_.joinable()) t_.join(); }
    ThreadGuard(const ThreadGuard &) = delete;
    ThreadGuard &operator=(const ThreadGuard &) = delete;
private:
    std::thread &t_;
};

void safe() {
    std::thread t(work, 7);
    ThreadGuard guard(t);          // if anything throws, the thread is still joined
    std::cout << "doing more work\n";
}   // guard's destructor joins here

Sharing data safely

When two threads touch the same variable without synchronisation you have a data race, and a data race is undefined behaviour — the result is not merely "wrong", the compiler may legally do anything. The example below increments a shared counter from four threads; without a lock the total comes out too small.

long total = 0;                    // SHARED — must be synchronised

void addUnsafe(int rounds) {
    for (int i = 0; i < rounds; ++i)
        ++total;                   // read-modify-write: three steps, not one
}
// Four threads each doing 100'000 rounds give a total well below 400'000,
// because updates overwrite each other.

Synchronization

To make a shared update safe you must make it atomic — one thread completes the whole read-modify-write before another begins. C++ gives you two families of tools: locks (mutexes) that guard a region of code, and atomics that make a single variable indivisible.

Mutexes and lock guards

A std::mutex is a lock. Wrap every access to the shared data in a std::lock_guard — an RAII object that unlocks in its destructor, so you cannot forget to release the lock even if an exception is thrown.

#include <mutex>

std::mutex m;
long total = 0;

void addSafe(int rounds) {
    for (int i = 0; i < rounds; ++i) {
        std::lock_guard<std::mutex> lock(m);   // locks here...
        ++total;                                // ...read-modify-write is now safe
    }                                           // ...and unlocks at the closing brace
}

For a critical section that must lock more than one mutex, use std::scoped_lock (C++17) — it locks all of them without deadlock. std::unique_lock is the flexible variant you need when you must unlock and re-lock manually, or hand the lock to a condition variable.

Atomics

When only a single variable needs protection, an std::atomic is usually faster and simpler than a mutex: its operations are indivisible and lock-free where the hardware supports it.

#include <atomic>

std::atomic<long> count{0};

void bump(int rounds) {
    for (int i = 0; i < rounds; ++i)
        count.fetch_add(1, std::memory_order_relaxed);  // one indivisible step
}

The default memory order (seq_cst) is the safest and easiest to reason about; relaxed is enough for a plain counter. Reach for the weaker orders only when you have measured a need and understand the visibility rules — they trade safety for speed.

ToolProtectsUse when
std::mutex + lock_guarda region of code / several variablesinvariants spanning more than one value
std::scoped_lockmany mutexes at onceyou must hold two locks together, deadlock-free
std::atomicone variablecounters, flags, single-word hand-off
std::condition_variablea wait/notify protocola thread waits for a state change

Tasks & futures

Raw threads force you to manage lifetimes by hand. The task-based model is higher level: you ask the standard library to run a function and hand you a std::future — a placeholder for a result that will be ready later. std::async is the simplest way to get one.

#include <future>
#include <numeric>

double heavy(int n) { /* ... expensive work ... */ return 0.0; }

int main() {
    std::future<double> job = std::async(std::launch::async, heavy, 1'000'000);
    // ... do unrelated work on the main thread while `job` runs ...
    double result = job.get();   // blocks until the task is finished, returns the value
}

get() may be called only once. Behind the scenes a future is fed by a promise or a packaged_task; those are the building blocks of thread pools and work queues, where many tasks are executed by a fixed set of worker threads.

Parallel patterns

The most common real pattern is "map": apply a function to every element of a range, in parallel. Split the range into chunks, hand each chunk to a worker, then join. Keeping each chunk to its own slice of the output means the workers never collide and no locking is needed.

#include <thread>
#include <vector>

// Out-of-place parallel map: out[i] = f(in[i]); workers write disjoint ranges.
void parallelMap(const std::vector<int> &in, std::vector<int> &out,
                 int (*f)(int), int workers) {
    out.resize(in.size());
    std::vector<std::thread> pool;
    const std::size_t n = in.size();
    for (int w = 0; w < workers; ++w) {
        std::size_t begin = n * w / workers;       // this worker's slice
        std::size_t end   = n * (w + 1) / workers;
        pool.emplace_back([&, begin, end] {
            for (std::size_t i = begin; i < end; ++i) out[i] = f(in[i]);
        });
    }
    for (auto &t : pool) t.join();
}

Because each thread writes a different region of out, there is no shared mutation and therefore no data race. C++17 goes further with execution policies: a single line requests a parallel standard algorithm without you writing any threading at all.

#include <algorithm>
#include <execution>

std::for_each(std::execution::par, data.begin(), data.end(), [](int &x) {
    x = x * x;          // may run on multiple threads
});

Pitfalls

Concurrency bugs are subtle because they hide in timing. Keep this list in mind and most of them will never reach production.

  • Data races are undefined behaviour. "It worked on my machine" is meaningless; use a lock or an atomic.
  • Deadlock. Two threads each waiting for the other's lock. Lock mutexes in one global order, or use std::scoped_lock.
  • Un-joined threads abort the program when a joinable std::thread is destroyed — always join() (or use a guard / std::jthread).
  • Over-locking destroys scalability. Shrink critical sections; prefer private data and merge at the end.
  • False sharing. Two threads writing different variables that sit on the same cache line still fight; separate them.
  • Threads are expensive. Do not spawn one per item — use a fixed pool sized to the core count.

Practice

  1. Run the data-race example in demo 11, then add the mutex and watch the total become exact.
  2. Convert the mutex counter to std::atomic<long> and compare the performance.
  3. Parallelise a sum over a large vector by giving each thread its own partial sum, then adding the partials at the end.
  4. Add a ThreadGuard to a function that throws, and prove the thread is still joined.
  5. Start three std::async tasks and collect their results with future::get().