Parallel & Multicore Processing
Approaches to Parallelism
Before writing any parallel code, name the kind of machine the work will run on. The approach you choose must match the hardware's memory model — this single decision shapes every API on this page.
Concurrency vs Parallelism
Concurrency is structure: a program that makes progress on several tasks at once, interleaving them so slow I/O does not block computation. Parallelism is execution: several tasks genuinely running at the same instant on different cores. A concurrent program can run on one core; a parallel program needs several. The two are not interchangeable — you can be concurrent without being parallel, and parallel without being concurrent.
Shared vs Distributed Memory
A shared-memory machine (one desktop, one virtual machine) gives every thread the same address space — fastest, but threads can step on each other. A distributed-memory cluster (many machines over a network) has no shared memory at all: processes must send messages. Threads, atomics, mutexes, OpenMP, and SIMD are shared-memory tools; MPI is the distributed-memory tool.
Threads — Multicore Execution
A thread is an independent execution flow that shares the process's memory. Two threads on two cores can genuinely run different code at the same instant — which is exactly why shared state must be protected (next chapter). The two main APIs are POSIX pthreads (the de facto standard on Linux and macOS) and the C11 <threads.h> interface (portable by the standard).
Figure 1 — one CPU-bound job split across cores: each thread runs a slice of the work (shared control-flow/HPC primitive).
pthreads — The De Facto Standard
POSIX threads are what production C actually uses on Linux and macOS. The pattern is always the same: create a thread with a function and an argument pointer, let it run, then join it to wait for completion.
#include <pthread.h> // POSIX threads
#include <stdio.h>
#define N 4
// worker: runs on its own thread; arg is a pointer to the thread's number
static void *worker(void *arg) {
int id = *(const int *)arg; // reinterpret the payload
printf("thread %d running on a core\n", id);
return NULL; // no result to return here
}
int main(void) {
pthread_t tid[N]; // one handle per thread
int ids[N] = {0, 1, 2, 3};
for (int i = 0; i < N; i++) {
// create: thread handle, NULL attrs, function, argument pointer
pthread_create(&tid[i], NULL, worker, &ids[i]);
}
for (int i = 0; i < N; i++) {
pthread_join(tid[i], NULL); // wait for thread i to finish
}
printf("all %d threads joined\n", N);
return 0;
}
// build: gcc -std=c11 -Wall -Wextra -Werror -pthread app.c -o app
The -pthread flag matters: it links the pthread library and defines the right feature macros. Without it, pthread_create becomes a link error on Linux at best, an undefined constant at worst.
C11 Threads — A Portable Alternative
The standard <threads.h> API mirrors pthreads with shorter names (thrd_create, thrd_join, thrd_t) and adds standard mutexes (mtx_t). It compiles anywhere the standard is implemented (glibc ≥ 2.28, musl, recent clang on BSDs) — but on MSVC and older glibc it is unavailable, which is why pthreads still dominates in the field.
#include <stdio.h>
#include <threads.h> // C11: thrd_t, thrd_create, thrd_join
static int greet(void *arg) { // signature required by thrd_create
printf("hello from %s\n", (const char *)arg);
return 0; // thrd_success
}
int main(void) {
thrd_t t;
if (thrd_create(&t, greet, "c11") != thrd_success) {
return 1; // thread could not be created
}
int result = 0;
thrd_join(t, &result); // wait; capture the thread's exit code
return 0;
}
Shared State & Synchronization
The moment two threads touch the same variable, the memory model of C applies rules with teeth. Read this chapter before parallelizing anything: the tools here are what keep shared state correct, and they are the ones most often misused.
Data Races
Two threads reading and writing the same memory at the same instant is a data race — undefined behavior in C, even for a single int. Compilers and CPUs may reorder and cache aggressively, and "just make it volatile" does not help (it fights only the compiler, not the hardware). The classic symptom: an increment that "loses" updates.
#include <pthread.h>
#include <stdio.h>
#define LOOP 100000
static int racy_count = 0; // NOT atomic: racing ++ is UB
static void *bump(void *arg) {
for (int i = 0; i < LOOP; i++) {
racy_count++; // BUG: read-modify-write interleaves badly
}
return NULL;
}
int main(void) {
pthread_t t1, t2;
pthread_create(&t1, NULL, bump, NULL);
pthread_create(&t2, NULL, bump, NULL);
pthread_join(t1, NULL);
pthread_join(t2, NULL);
printf("racing counter = %d (expected %d)\n", racy_count, 2 * LOOP);
return 0;
}
Run it: the result is usually less than 200000, sometimes far less. Under ThreadSanitizer (-fsanitize=thread) the race is reported line by line with both participants. There is no compiler fix for a race — only restructuring.
Atomics
Atomic types from <stdatomic.h> perform each operation as one indivisible step the hardware guarantees. For counters and flags, atomics are faster than locks and lock-free by design.
#include <pthread.h>
#include <stdatomic.h> // atomic_int, atomic_fetch_add
#include <stdio.h>
#define LOOP 100000
static atomic_int safe_count = 0; // atomic: every op is indivisible
static void *bump(void *arg) {
for (int i = 0; i < LOOP; i++) {
atomic_fetch_add(&safe_count, 1); // one atomic read-modify-write
}
return NULL;
}
int main(void) {
pthread_t t1, t2;
pthread_create(&t1, NULL, bump, NULL);
pthread_create(&t2, NULL, bump, NULL);
pthread_join(t1, NULL);
pthread_join(t2, NULL);
printf("atomic counter = %d (exactly %d)\n",
atomic_load(&safe_count), 2 * LOOP);
return 0;
}
The default memory order is memory_order_seq_cst — the strongest, easiest-to-reason guarantee; relax it only after you understand the weaker orders. If the shared region is bigger than one integer, atomics are not enough — that is what mutexes are for.
Mutexes & Locks
A mutex (mutual exclusion) lets only one thread inside a critical section at a time, turning a sequence of statements into an indivisible region regardless of size. Atomics scale better for tiny counters; mutexes are correct for anything more.
#include <pthread.h>
#include <stdio.h>
static pthread_mutex_t gate = PTHREAD_MUTEX_INITIALIZER; // static init
static int account = 1000; // the safeguarded shared balance
static void *withdraw(void *arg) {
int amount = *(const int *)arg;
for (int i = 0; i < 10; i++) {
pthread_mutex_lock(&gate); // enter the critical section
if (account >= amount) {
account -= amount; // check-then-act, now atomic
}
pthread_mutex_unlock(&gate); // leave — another thread may enter
}
return NULL;
}
int main(void) {
int five = 5;
pthread_t a, b;
pthread_create(&a, NULL, withdraw, &five);
pthread_create(&b, NULL, withdraw, &five);
pthread_join(a, NULL);
pthread_join(b, NULL);
pthread_mutex_destroy(&gate); // release the OS resource
printf("balance = %d\n", account); // never negative: check-then-act held
return 0;
}
Every lock must have a matching unlock on every path — an early return inside a locked section is a classic deadlock. When a thread must wait for an event or condition, combine a mutex with a condition variable (pthread_cond_t); when only one thread should ever run a block, use a once flag (pthread_once).
Data Parallelism on the CPU
Not all parallelism needs explicit threads. Pure data transform loops — "apply this formula to every element" — can be parallelized without a thread API in sight: OpenMP shares the loop across cores, and the CPU's SIMD units process several values per instruction on a single core.
OpenMP — Pragmas for Multicore
OpenMP turns loops into parallel code with a single pragma; the compiler generates the thread management for you. It is the fastest route to a multicore speedup for CPU-bound loops over independent data, and it does not change your program's structure. Compile with -fopenmp (GCC and Clang).
#include <omp.h> // omp_get_num_threads
#include <stdio.h>
int main(void) {
long double sum = 0.0L;
// each iteration runs on whatever core; reduction(+:sum) merges
// the per-thread partial sums into one final value, race-free
#pragma omp parallel for reduction(+:sum)
for (long i = 1; i <= 1000000; i++) {
sum += 1.0L / (long double)i; // harmonic series, fine to split
}
printf("harmonic sum = %.6Lf using %d threads\n",
sum, omp_get_num_threads());
return 0;
}
// build: gcc -std=c11 -Wall -Wextra -fopenmp app.c -o app
The reduction clause is what keeps this correct: each thread keeps a private partial sum, and the runtime combines them at the end — no shared-variable race exists at all. That one clause removes the entire synchronization chapter from this code path.
SIMD & Auto-Vectorization
Inside a single core, modern CPUs process several values per instruction — SIMD (single instruction, multiple data). The free first step: compile with -O2 -march=native and promise no aliasing with restrict (seen on the advanced page), and GCC/Clang fuse simple loops into SIMD instructions. When the compiler cannot, write intrinsics directly.
#include <immintrin.h> // AVX2 intrinsics
#include <stdio.h>
int main(void) {
// eight floats per AVX lane: one instruction adds 8 values at once
__m256 a = _mm256_setr_ps(1, 2, 3, 4, 5, 6, 7, 8);
__m256 b = _mm256_setr_ps(8, 7, 6, 5, 4, 3, 2, 1);
__m256 c = _mm256_add_ps(a, b); // packed add: 8 results in one op
float out[8];
_mm256_storeu_ps(out, c); // write the 8 lanes to memory
printf("c[0] = %g, c[7] = %g\n", out[0], out[7]); // 9 and 9
return 0;
}
// build: gcc -std=c11 -march=native -O2 app.c -o app
Intrinsics grow with the ISA (AVX, AVX2, AVX-512 on x86; NEON/SVE on ARM) — each has its own header, so intrinsic code is less portable than auto-vectorized code. Start with the compiler route; reach for intrinsics only on measured hot paths.
Distributed Parallelism
When the problem outgrows one machine, threads no longer work — there is no shared memory across a network. The tool changes from threads to messages.
MPI — Message Passing
MPI (Message Passing Interface) treats each machine as a separate process with its own memory, exchanging data with explicit send/receive calls. It is the standard for scientific computing on clusters, and C is one of its native languages.
# typical MPI build & run (OpenMPI or MPICH installed)
mpicc -std=c11 -Wall -Wextra app.c -o app
mpiexec -np 4 ./app # launch 4 cooperating processes
MPI code lives between MPI_Init and MPI_Finalize, introspects with MPI_Comm_rank/MPI_Comm_size, and moves data with collectives such as MPI_Reduce and MPI_Bcast. The mental model is entirely different from threads: you write a single program, multiple data (SPMD) kernel that every rank runs. Need depth? The dedicated HPC roadmap covers the full stack.
Choosing an Approach
The approaches here are tools, not religions. Match the tool to the bottleneck:
| Scenario | Approach | Tool | Build flag |
|---|---|---|---|
| 2–64 independent big tasks, one machine | Threads | pthreads / <threads.h> | -pthread |
| Shared counters or flags | Atomics | <stdatomic.h> | header-only |
| Larger shared critical sections | Mutexes | pthread_mutex_t | -pthread |
| CPU-heavy loops over independent data | OpenMP | #pragma omp parallel for | -fopenmp |
| Dense packed math on one core | SIMD | auto-vectorize / intrinsics | -O2 -march=native |
| Many machines over a network | MPI | OpenMPI / MPICH | mpicc |
time or clock_gettime before and after. Parallelism you cannot measure is marketing.
Next lab: GPU computing — moving the hot loops to a thousand-core co-processor. After that, ABI & assembly shows what the machine really does when your functions call each other.