Parallel Computing

Fortran is a parallel language by design: do concurrent and coarrays are part of the standard, and OpenMP, MPI, OpenACC and CUDA Fortran are first-class companions. The first decision is not which API to type but which hardware you are aiming at. This lesson maps the landscape and teaches the shared-memory model in depth; the next lessons drill into coarrays, MPI and GPUs.

The parallel landscape

Parallel programming splits along one axis: whether the threads of execution share one memory or not.

  • Shared memory (one CPU, many cores): threads see the same variables. OpenMP and do concurrent live here. Low latency, but races and cache coherence are the enemies.
  • Distributed memory (many nodes): processes own private memory and communicate by message. MPI lives here; coarrays express the same idea with one-sided reads. Scales to the top-500 machines.
  • Accelerators (GPU): a device with thousands of simple cores and its own memory, driven by offload directives or CUDA. OpenACC, OpenMP target, CUDA Fortran.

Reality is a hybrid stack: MPI between nodes, OpenMP inside a node, GPU offload for the hot kernel. The architecture is decided first, then the APIs follow — never the reverse.

OpenMP: the shared-memory workhorse

OpenMP is a set of compiler directives (!$omp ...) that parallelize existing code incrementally: the compiler generates the threads, the distribution of loop iterations, and the synchronization. It is specified by the OpenMP ARB for C, C++ and Fortran, supports every mainstream compiler (-fopenmp), and its directives are ordinary comments for compilers without support — the serial version still compiles and runs.

OpenMP fork-join execution model

Fig. 1 — One master thread forks a team, parallel region runs, barrier, join.

Parallel loops and reductions

The workhorse directive is !$omp parallel do, which distributes loop iterations across the team. Every iteration must be independent — any shared value written by one iteration and read by another is a race. The reduction(+:sum) clause handles the single most common pattern, a summation, correctly: each thread accumulates a private copy and the copies are combined at the join.

program pi_openmp
  use omp_lib                     ! the OpenMP module (with -fopenmp)
  implicit none
  integer, parameter :: n = 100000000
  real(8) :: sum, x, h
  integer :: i, nthreads
  h = 1.0_8 / n
  sum = 0.0_8
  !$omp parallel do reduction(+:sum) private(x)   ! race-free summation
  do i = 1, n
     x = (i - 0.5_8) * h
     sum = sum + 4.0_8 / (1.0_8 + x * x)
  end do
  !$omp end parallel do
  print '(a,f12.9)', 'pi ~', h * sum
  nthreads = omp_get_max_threads()
  print '(a,i0,a)', 'ran with ', nthreads, ' threads'
end program pi_openmp

Compile with gfortran -fopenmp -O3 pi_openmp.f90 and run with OMP_NUM_THREADS=8 ./pi_openmp (or set OMP_NUM_THREADS=8 on Windows). Compare wall time against OMP_NUM_THREADS=1: the loop is embarrassingly parallel, so the speedup should approach the thread count. The private(x) clause gives every thread its own x; without it the threads would fight over one shared variable and the answer would be wrong (or NaN).

Thread safety and the race class

The compiler cannot know that two iterations are independent — it only parallelizes what you tell it to, and correctness is your contract. The race class to internalize: a read-modify-write on a shared variable (sum = sum + v(i)) is illegal without a reduction. firstprivate copies a value into each thread at entry; critical serializes a whole block (use sparingly: it kills the parallelism it protects); atomic protects a single statement. For arrays, the gold rule from the arrays lesson — each thread writes only its own slice — turns most loops trivially safe.

program histogram
  implicit none
  integer, parameter :: nbins = 100, n = 1000000
  real(8) :: v(n)
  integer :: hist(nbins), i, b
  hist = 0
  call random_number(v)
  !$omp parallel do private(b) reduction(+:hist)
  do i = 1, n
     b = 1 + int(v(i) * nbins)      ! each iteration touches ONE bin
     hist(b) = hist(b) + 1          ! but bins are shared → reduce
  end do
  !$omp end parallel do
  print '(a,i0)', 'total count = ', sum(hist)
end program histogram

DO CONCURRENT vs OpenMP

do concurrent is the standard's own parallel loop: it declares iteration independence (control lesson) and lets each compiler choose threads, SIMD width, or serial order. OpenMP is imperative: you name the parallel region and the distribution. In practice, do concurrent is cleaner for new code and gfortran/ifx can map it onto OpenMP under the hood; OpenMP remains essential for reductions before F2023 support, for explicit thread management, and for the whole family of clauses (schedule, nowait, ordered, tasks, target offload). A pragmatic split: do concurrent for loops, OpenMP when you need the control.

Choosing a model: the decision table

Your hardwareFirst choiceWhyLesson
one multicore nodeOpenMP / do concurrentlowest latency, incrementally addedthis lesson
one node, GPU attachedOpenACC or OpenMP targetportable across vendors and CPUsGPU lesson
many nodes / clusterMPIthe only portable distributed standardMPI lesson
HPC cluster, Fortran-nativecoarraysone-sided communication in the standard itselfCoarrays lesson

Hybrid is the norm at the top: MPI between nodes, OpenMP inside each node, a GPU offload for the densest kernel. The three following lessons build each layer with working code and profiling advice.