Parallel Computing
do concurrent and coarrays are part of the standard, and OpenMP, MPI, OpenACC and CUDA Fortran are first-class companions. The first decision is not which API to type but which hardware you are aiming at. This lesson maps the landscape and teaches the shared-memory model in depth; the next lessons drill into coarrays, MPI and GPUs.
The parallel landscape
Parallel programming splits along one axis: whether the threads of execution share one memory or not.
- Shared memory (one CPU, many cores): threads see the same variables. OpenMP and
do concurrentlive here. Low latency, but races and cache coherence are the enemies. - Distributed memory (many nodes): processes own private memory and communicate by message. MPI lives here; coarrays express the same idea with one-sided reads. Scales to the top-500 machines.
- Accelerators (GPU): a device with thousands of simple cores and its own memory, driven by offload directives or CUDA. OpenACC, OpenMP target, CUDA Fortran.
Reality is a hybrid stack: MPI between nodes, OpenMP inside a node, GPU offload for the hot kernel. The architecture is decided first, then the APIs follow — never the reverse.
OpenMP: the shared-memory workhorse
OpenMP is a set of compiler directives (!$omp ...) that parallelize existing code incrementally: the compiler generates the threads, the distribution of loop iterations, and the synchronization. It is specified by the OpenMP ARB for C, C++ and Fortran, supports every mainstream compiler (-fopenmp), and its directives are ordinary comments for compilers without support — the serial version still compiles and runs.
Fig. 1 — One master thread forks a team, parallel region runs, barrier, join.
Parallel loops and reductions
The workhorse directive is !$omp parallel do, which distributes loop iterations across the team. Every iteration must be independent — any shared value written by one iteration and read by another is a race. The reduction(+:sum) clause handles the single most common pattern, a summation, correctly: each thread accumulates a private copy and the copies are combined at the join.
program pi_openmp
use omp_lib ! the OpenMP module (with -fopenmp)
implicit none
integer, parameter :: n = 100000000
real(8) :: sum, x, h
integer :: i, nthreads
h = 1.0_8 / n
sum = 0.0_8
!$omp parallel do reduction(+:sum) private(x) ! race-free summation
do i = 1, n
x = (i - 0.5_8) * h
sum = sum + 4.0_8 / (1.0_8 + x * x)
end do
!$omp end parallel do
print '(a,f12.9)', 'pi ~', h * sum
nthreads = omp_get_max_threads()
print '(a,i0,a)', 'ran with ', nthreads, ' threads'
end program pi_openmp
Compile with gfortran -fopenmp -O3 pi_openmp.f90 and run with OMP_NUM_THREADS=8 ./pi_openmp (or set OMP_NUM_THREADS=8 on Windows). Compare wall time against OMP_NUM_THREADS=1: the loop is embarrassingly parallel, so the speedup should approach the thread count. The private(x) clause gives every thread its own x; without it the threads would fight over one shared variable and the answer would be wrong (or NaN).
Thread safety and the race class
The compiler cannot know that two iterations are independent — it only parallelizes what you tell it to, and correctness is your contract. The race class to internalize: a read-modify-write on a shared variable (sum = sum + v(i)) is illegal without a reduction. firstprivate copies a value into each thread at entry; critical serializes a whole block (use sparingly: it kills the parallelism it protects); atomic protects a single statement. For arrays, the gold rule from the arrays lesson — each thread writes only its own slice — turns most loops trivially safe.
program histogram
implicit none
integer, parameter :: nbins = 100, n = 1000000
real(8) :: v(n)
integer :: hist(nbins), i, b
hist = 0
call random_number(v)
!$omp parallel do private(b) reduction(+:hist)
do i = 1, n
b = 1 + int(v(i) * nbins) ! each iteration touches ONE bin
hist(b) = hist(b) + 1 ! but bins are shared → reduce
end do
!$omp end parallel do
print '(a,i0)', 'total count = ', sum(hist)
end program histogram
DO CONCURRENT vs OpenMP
do concurrent is the standard's own parallel loop: it declares iteration independence (control lesson) and lets each compiler choose threads, SIMD width, or serial order. OpenMP is imperative: you name the parallel region and the distribution. In practice, do concurrent is cleaner for new code and gfortran/ifx can map it onto OpenMP under the hood; OpenMP remains essential for reductions before F2023 support, for explicit thread management, and for the whole family of clauses (schedule, nowait, ordered, tasks, target offload). A pragmatic split: do concurrent for loops, OpenMP when you need the control.
Choosing a model: the decision table
| Your hardware | First choice | Why | Lesson |
|---|---|---|---|
| one multicore node | OpenMP / do concurrent | lowest latency, incrementally added | this lesson |
| one node, GPU attached | OpenACC or OpenMP target | portable across vendors and CPUs | GPU lesson |
| many nodes / cluster | MPI | the only portable distributed standard | MPI lesson |
| HPC cluster, Fortran-native | coarrays | one-sided communication in the standard itself | Coarrays lesson |
Hybrid is the norm at the top: MPI between nodes, OpenMP inside each node, a GPU offload for the densest kernel. The three following lessons build each layer with working code and profiling advice.