Parallel & Multicore Processing

Every modern CPU chip has several cores, and C is the language that lets you use them directly. This lab covers CPU parallelism specifically: threads for a few big tasks, atomics and mutexes for safe shared state, OpenMP for simple loop speedups, SIMD for data-parallel math on a single core, and MPI for many machines. GPU programming — the thousand-core co-processor — is its own GPU lab that follows.

Approaches to Parallelism

Before writing any parallel code, name the kind of machine the work will run on. The approach you choose must match the hardware's memory model — this single decision shapes every API on this page.

Concurrency vs Parallelism

Concurrency is structure: a program that makes progress on several tasks at once, interleaving them so slow I/O does not block computation. Parallelism is execution: several tasks genuinely running at the same instant on different cores. A concurrent program can run on one core; a parallel program needs several. The two are not interchangeable — you can be concurrent without being parallel, and parallel without being concurrent.

Shared vs Distributed Memory

A shared-memory machine (one desktop, one virtual machine) gives every thread the same address space — fastest, but threads can step on each other. A distributed-memory cluster (many machines over a network) has no shared memory at all: processes must send messages. Threads, atomics, mutexes, OpenMP, and SIMD are shared-memory tools; MPI is the distributed-memory tool.

Threads — Multicore Execution

A thread is an independent execution flow that shares the process's memory. Two threads on two cores can genuinely run different code at the same instant — which is exactly why shared state must be protected (next chapter). The two main APIs are POSIX pthreads (the de facto standard on Linux and macOS) and the C11 <threads.h> interface (portable by the standard).

One job split across several cores

Figure 1 — one CPU-bound job split across cores: each thread runs a slice of the work (shared control-flow/HPC primitive).

pthreads — The De Facto Standard

POSIX threads are what production C actually uses on Linux and macOS. The pattern is always the same: create a thread with a function and an argument pointer, let it run, then join it to wait for completion.

#include <pthread.h>   // POSIX threads
#include <stdio.h>

#define N 4

// worker: runs on its own thread; arg is a pointer to the thread's number
static void *worker(void *arg) {
    int id = *(const int *)arg;              // reinterpret the payload
    printf("thread %d running on a core\n", id);
    return NULL;                             // no result to return here
}

int main(void) {
    pthread_t tid[N];                        // one handle per thread
    int ids[N] = {0, 1, 2, 3};
    for (int i = 0; i < N; i++) {
        // create: thread handle, NULL attrs, function, argument pointer
        pthread_create(&tid[i], NULL, worker, &ids[i]);
    }
    for (int i = 0; i < N; i++) {
        pthread_join(tid[i], NULL);          // wait for thread i to finish
    }
    printf("all %d threads joined\n", N);
    return 0;
}
// build: gcc -std=c11 -Wall -Wextra -Werror -pthread app.c -o app

The -pthread flag matters: it links the pthread library and defines the right feature macros. Without it, pthread_create becomes a link error on Linux at best, an undefined constant at worst.

C11 Threads — A Portable Alternative

The standard <threads.h> API mirrors pthreads with shorter names (thrd_create, thrd_join, thrd_t) and adds standard mutexes (mtx_t). It compiles anywhere the standard is implemented (glibc ≥ 2.28, musl, recent clang on BSDs) — but on MSVC and older glibc it is unavailable, which is why pthreads still dominates in the field.

#include <stdio.h>
#include <threads.h>   // C11: thrd_t, thrd_create, thrd_join

static int greet(void *arg) {                // signature required by thrd_create
    printf("hello from %s\n", (const char *)arg);
    return 0;                                // thrd_success
}

int main(void) {
    thrd_t t;
    if (thrd_create(&t, greet, "c11") != thrd_success) {
        return 1;                            // thread could not be created
    }
    int result = 0;
    thrd_join(t, &result);                   // wait; capture the thread's exit code
    return 0;
}

Shared State & Synchronization

The moment two threads touch the same variable, the memory model of C applies rules with teeth. Read this chapter before parallelizing anything: the tools here are what keep shared state correct, and they are the ones most often misused.

Data Races

Two threads reading and writing the same memory at the same instant is a data race — undefined behavior in C, even for a single int. Compilers and CPUs may reorder and cache aggressively, and "just make it volatile" does not help (it fights only the compiler, not the hardware). The classic symptom: an increment that "loses" updates.

#include <pthread.h>
#include <stdio.h>

#define LOOP 100000

static int racy_count = 0;   // NOT atomic: racing ++ is UB

static void *bump(void *arg) {
    for (int i = 0; i < LOOP; i++) {
        racy_count++;            // BUG: read-modify-write interleaves badly
    }
    return NULL;
}

int main(void) {
    pthread_t t1, t2;
    pthread_create(&t1, NULL, bump, NULL);
    pthread_create(&t2, NULL, bump, NULL);
    pthread_join(t1, NULL);
    pthread_join(t2, NULL);
    printf("racing counter = %d (expected %d)\n", racy_count, 2 * LOOP);
    return 0;
}

Run it: the result is usually less than 200000, sometimes far less. Under ThreadSanitizer (-fsanitize=thread) the race is reported line by line with both participants. There is no compiler fix for a race — only restructuring.

Atomics

Atomic types from <stdatomic.h> perform each operation as one indivisible step the hardware guarantees. For counters and flags, atomics are faster than locks and lock-free by design.

#include <pthread.h>
#include <stdatomic.h>   // atomic_int, atomic_fetch_add
#include <stdio.h>

#define LOOP 100000

static atomic_int safe_count = 0;   // atomic: every op is indivisible

static void *bump(void *arg) {
    for (int i = 0; i < LOOP; i++) {
        atomic_fetch_add(&safe_count, 1);   // one atomic read-modify-write
    }
    return NULL;
}

int main(void) {
    pthread_t t1, t2;
    pthread_create(&t1, NULL, bump, NULL);
    pthread_create(&t2, NULL, bump, NULL);
    pthread_join(t1, NULL);
    pthread_join(t2, NULL);
    printf("atomic counter = %d (exactly %d)\n",
           atomic_load(&safe_count), 2 * LOOP);
    return 0;
}

The default memory order is memory_order_seq_cst — the strongest, easiest-to-reason guarantee; relax it only after you understand the weaker orders. If the shared region is bigger than one integer, atomics are not enough — that is what mutexes are for.

Mutexes & Locks

A mutex (mutual exclusion) lets only one thread inside a critical section at a time, turning a sequence of statements into an indivisible region regardless of size. Atomics scale better for tiny counters; mutexes are correct for anything more.

#include <pthread.h>
#include <stdio.h>

static pthread_mutex_t gate = PTHREAD_MUTEX_INITIALIZER;  // static init
static int account = 1000;          // the safeguarded shared balance

static void *withdraw(void *arg) {
    int amount = *(const int *)arg;
    for (int i = 0; i < 10; i++) {
        pthread_mutex_lock(&gate);        // enter the critical section
        if (account >= amount) {
            account -= amount;            // check-then-act, now atomic
        }
        pthread_mutex_unlock(&gate);      // leave — another thread may enter
    }
    return NULL;
}

int main(void) {
    int five = 5;
    pthread_t a, b;
    pthread_create(&a, NULL, withdraw, &five);
    pthread_create(&b, NULL, withdraw, &five);
    pthread_join(a, NULL);
    pthread_join(b, NULL);
    pthread_mutex_destroy(&gate);         // release the OS resource
    printf("balance = %d\n", account);    // never negative: check-then-act held
    return 0;
}

Every lock must have a matching unlock on every path — an early return inside a locked section is a classic deadlock. When a thread must wait for an event or condition, combine a mutex with a condition variable (pthread_cond_t); when only one thread should ever run a block, use a once flag (pthread_once).

Data Parallelism on the CPU

Not all parallelism needs explicit threads. Pure data transform loops — "apply this formula to every element" — can be parallelized without a thread API in sight: OpenMP shares the loop across cores, and the CPU's SIMD units process several values per instruction on a single core.

OpenMP — Pragmas for Multicore

OpenMP turns loops into parallel code with a single pragma; the compiler generates the thread management for you. It is the fastest route to a multicore speedup for CPU-bound loops over independent data, and it does not change your program's structure. Compile with -fopenmp (GCC and Clang).

#include <omp.h>       // omp_get_num_threads
#include <stdio.h>

int main(void) {
    long double sum = 0.0L;
    // each iteration runs on whatever core; reduction(+:sum) merges
    // the per-thread partial sums into one final value, race-free
    #pragma omp parallel for reduction(+:sum)
    for (long i = 1; i <= 1000000; i++) {
        sum += 1.0L / (long double)i;       // harmonic series, fine to split
    }
    printf("harmonic sum = %.6Lf using %d threads\n",
           sum, omp_get_num_threads());
    return 0;
}
// build: gcc -std=c11 -Wall -Wextra -fopenmp app.c -o app

The reduction clause is what keeps this correct: each thread keeps a private partial sum, and the runtime combines them at the end — no shared-variable race exists at all. That one clause removes the entire synchronization chapter from this code path.

SIMD & Auto-Vectorization

Inside a single core, modern CPUs process several values per instruction — SIMD (single instruction, multiple data). The free first step: compile with -O2 -march=native and promise no aliasing with restrict (seen on the advanced page), and GCC/Clang fuse simple loops into SIMD instructions. When the compiler cannot, write intrinsics directly.

#include <immintrin.h>  // AVX2 intrinsics
#include <stdio.h>

int main(void) {
    // eight floats per AVX lane: one instruction adds 8 values at once
    __m256 a = _mm256_setr_ps(1, 2, 3, 4, 5, 6, 7, 8);
    __m256 b = _mm256_setr_ps(8, 7, 6, 5, 4, 3, 2, 1);
    __m256 c = _mm256_add_ps(a, b);       // packed add: 8 results in one op
    float out[8];
    _mm256_storeu_ps(out, c);             // write the 8 lanes to memory
    printf("c[0] = %g, c[7] = %g\n", out[0], out[7]);   // 9 and 9
    return 0;
}
// build: gcc -std=c11 -march=native -O2 app.c -o app

Intrinsics grow with the ISA (AVX, AVX2, AVX-512 on x86; NEON/SVE on ARM) — each has its own header, so intrinsic code is less portable than auto-vectorized code. Start with the compiler route; reach for intrinsics only on measured hot paths.

Distributed Parallelism

When the problem outgrows one machine, threads no longer work — there is no shared memory across a network. The tool changes from threads to messages.

MPI — Message Passing

MPI (Message Passing Interface) treats each machine as a separate process with its own memory, exchanging data with explicit send/receive calls. It is the standard for scientific computing on clusters, and C is one of its native languages.

# typical MPI build & run (OpenMPI or MPICH installed)
mpicc -std=c11 -Wall -Wextra app.c -o app
mpiexec -np 4 ./app        # launch 4 cooperating processes

MPI code lives between MPI_Init and MPI_Finalize, introspects with MPI_Comm_rank/MPI_Comm_size, and moves data with collectives such as MPI_Reduce and MPI_Bcast. The mental model is entirely different from threads: you write a single program, multiple data (SPMD) kernel that every rank runs. Need depth? The dedicated HPC roadmap covers the full stack.

Choosing an Approach

The approaches here are tools, not religions. Match the tool to the bottleneck:

ScenarioApproachToolBuild flag
2–64 independent big tasks, one machineThreadspthreads / <threads.h>-pthread
Shared counters or flagsAtomics<stdatomic.h>header-only
Larger shared critical sectionsMutexespthread_mutex_t-pthread
CPU-heavy loops over independent dataOpenMP#pragma omp parallel for-fopenmp
Dense packed math on one coreSIMDauto-vectorize / intrinsics-O2 -march=native
Many machines over a networkMPIOpenMPI / MPICHmpicc
Homework to apply this lab: take the word-frequency counter from the study projects and speed it up three ways — OpenMP over file chunks, an atomic counter for the totals, then measure with time or clock_gettime before and after. Parallelism you cannot measure is marketing.

Next lab: GPU computing — moving the hot loops to a thousand-core co-processor. After that, ABI & assembly shows what the machine really does when your functions call each other.