Fortran Performance
The optimization ladder
Compilers offer escalating effort levels. The default (-O0) keeps variables in memory for debuggability and is slow; -O2 applies the safe classical optimizations; -O3 adds vectorization and inlining that change floating-point results slightly — a 1e-15 reshuffle, fine for most science, fatal for strict reproducibility; -Ofast additionally breaks strict IEEE compliance. The ladder for production: never below -O2, -O3 for compute kernels after a correctness pass at -O0 -fcheck=bounds.
| Flag | What the compiler does | When to use |
|---|---|---|
-O0 | no optimization, variables in memory | debugging, sanitizer runs |
-O2 | classical optimizations, no precision change | production default |
-O3 | + vectorization, inlining, loop transforms | compute kernels, after tests pass |
-march=native | target the exact CPU instruction set | own machine only (ships: -march=x86-64-v3) |
-flto | link-time optimization across files | release builds |
-fopenmp/-fcoarray/-fopenacc | enable the parallel models | parallel lessons below |
How memory defeats naive code
RAM latency is roughly 200 cycles; a well-ordered do loop hides it by reading sequential cache lines — the prefetcher's favourite pattern — while a stride-4K walk misses every time and runs 10–50× slower. Combined with the column-major layout from the arrays lesson, this yields the two rules that dominate real kernels: walk the first index innermost and keep one element per loop iteration contiguous.
subroutine stencil(u, v, n)
! 2-D diffusion step: v(i,j) = average of the four neighbours of u.
integer, intent(in) :: n
real(8), intent(in) :: u(n, n)
real(8), intent(inout) :: v(n, n)
integer :: i, j
do j = 2, n - 1 ! outer: columns
do i = 2, n - 1 ! inner: rows — walks memory in order
v(i, j) = 0.25_8 * (u(i-1, j) + u(i+1, j) + u(i, j-1) + u(i, j+1))
end do
end do
end subroutine stencil
Vectorization: the second lever
Vectorization makes one instruction process several data elements (AVX-512: 8 doubles at once). The compiler vectorizes loops that are: unit-stride in the innermost index, free of cross-iteration dependencies, and free of function calls it cannot inline. The -O3 flag enables it; -fopt-info-vec-optimized (gfortran) prints which loops got vectorized — the fastest honesty check in the toolkit. A stencil like the one above vectorizes cleanly; a loop that calls sqrt also does (intrinsic-aware), but a loop calling your module function only after elemental and inlining.
Profile first: Amdahl and the hot spot
Parallelism obeys Amdahl’s law: the unparallelizable fraction sets the ceiling. It makes no sense to parallelize a loop that runs 1% of the time; the first tool is the profiler, not the directive.
Fig. 2 — Even 5% serial code caps the speedup near 20×; find the hot spot first.
The workflow is three steps and a loop: (1) compile with -pg and run to record call counts (gprof: gprof ./app gmon.out); (2) read the flat profile — the top function is where the seconds go; (3) optimize exactly that function: argue first (is the algorithm quadratic?), then flags, then memory order, then parallelize. Re-profile after every change; without measurement you are guessing, and guessing in performance work is expensive.
Correctness gates before speed
Optimization invalidates nothing you learned about safety: debug profile first (-O0 -fcheck=bounds -g -fbacktrace), release after (-O3 -flto -march=native). For numerics, insert a reproducibility guard — run the release build against -ffloat-store or a reference configuration and compare with a tolerance. The IEEE NaN traps from the errors lesson catch the silent corruption that optimization sometimes uncovers. Speed without a correctness gate is a faster wrong answer.
-O3 vectorize it, then hand it to do concurrent or OpenMP — measuring at every step.