Fortran Performance

Fortran's reputation rests on one promise: fast arithmetic over big arrays. That promise is not automatic — the compiler must be told how hard to try, the memory layout must match how loops walk it, and the bottleneck must be found before it is fixed. This lesson is the systematic performance workflow: measure, set flags, fix memory access, parallelize.

The optimization ladder

Compilers offer escalating effort levels. The default (-O0) keeps variables in memory for debuggability and is slow; -O2 applies the safe classical optimizations; -O3 adds vectorization and inlining that change floating-point results slightly — a 1e-15 reshuffle, fine for most science, fatal for strict reproducibility; -Ofast additionally breaks strict IEEE compliance. The ladder for production: never below -O2, -O3 for compute kernels after a correctness pass at -O0 -fcheck=bounds.

FlagWhat the compiler doesWhen to use
-O0no optimization, variables in memorydebugging, sanitizer runs
-O2classical optimizations, no precision changeproduction default
-O3+ vectorization, inlining, loop transformscompute kernels, after tests pass
-march=nativetarget the exact CPU instruction setown machine only (ships: -march=x86-64-v3)
-fltolink-time optimization across filesrelease builds
-fopenmp/-fcoarray/-fopenaccenable the parallel modelsparallel lessons below

How memory defeats naive code

RAM latency is roughly 200 cycles; a well-ordered do loop hides it by reading sequential cache lines — the prefetcher's favourite pattern — while a stride-4K walk misses every time and runs 10–50× slower. Combined with the column-major layout from the arrays lesson, this yields the two rules that dominate real kernels: walk the first index innermost and keep one element per loop iteration contiguous.

subroutine stencil(u, v, n)
  ! 2-D diffusion step: v(i,j) = average of the four neighbours of u.
  integer, intent(in) :: n
  real(8), intent(in) :: u(n, n)
  real(8), intent(inout) :: v(n, n)
  integer :: i, j
  do j = 2, n - 1            ! outer: columns
     do i = 2, n - 1         ! inner: rows — walks memory in order
        v(i, j) = 0.25_8 * (u(i-1, j) + u(i+1, j) + u(i, j-1) + u(i, j+1))
     end do
  end do
end subroutine stencil

Vectorization: the second lever

Vectorization makes one instruction process several data elements (AVX-512: 8 doubles at once). The compiler vectorizes loops that are: unit-stride in the innermost index, free of cross-iteration dependencies, and free of function calls it cannot inline. The -O3 flag enables it; -fopt-info-vec-optimized (gfortran) prints which loops got vectorized — the fastest honesty check in the toolkit. A stencil like the one above vectorizes cleanly; a loop that calls sqrt also does (intrinsic-aware), but a loop calling your module function only after elemental and inlining.

Profile first: Amdahl and the hot spot

Parallelism obeys Amdahl’s law: the unparallelizable fraction sets the ceiling. It makes no sense to parallelize a loop that runs 1% of the time; the first tool is the profiler, not the directive.

Amdahl's law speedup curves

Fig. 2 — Even 5% serial code caps the speedup near 20×; find the hot spot first.

The workflow is three steps and a loop: (1) compile with -pg and run to record call counts (gprof: gprof ./app gmon.out); (2) read the flat profile — the top function is where the seconds go; (3) optimize exactly that function: argue first (is the algorithm quadratic?), then flags, then memory order, then parallelize. Re-profile after every change; without measurement you are guessing, and guessing in performance work is expensive.

Correctness gates before speed

Optimization invalidates nothing you learned about safety: debug profile first (-O0 -fcheck=bounds -g -fbacktrace), release after (-O3 -flto -march=native). For numerics, insert a reproducibility guard — run the release build against -ffloat-store or a reference configuration and compare with a tolerance. The IEEE NaN traps from the errors lesson catch the silent corruption that optimization sometimes uncovers. Speed without a correctness gate is a faster wrong answer.

The one-sentence workflow: identify the hot loop with a profiler, keep its innermost index contiguous, let -O3 vectorize it, then hand it to do concurrent or OpenMP — measuring at every step.