Coarray Fortran

Coarrays (Fortran 2008/2018) put distributed-memory parallelism inside the language: one program runs as N replicated images, each owns private data plus coarrays, and any image reads another's coarray with square-bracket syntax. The compiler and runtime own the communication.

Images and cosubscripts

Launch ./app and num_images() tells you how many copies the runtime created (1 for a normal run; N with a coarray-aware scheduler). this_image() is your rank. A coarray is any variable with a cosubscript dimension — the trailing [*] — existing independently on every image. The square bracket selects which image's copy you mean, and reading it is one-sided: the owner is neither told nor paused:

program images
  implicit none
  integer :: me
  real :: buf[*]                    ! one copy per image, cosubscript address
  me = this_image()                 ! which image am I? 1..num_images()
  buf = real(me)                    ! write MY OWN copy (plain assignment)
  sync all                          ! barrier: now every buf is published
  if (me == 1) then
     print '(a,i0,a,f4.1)', 'image 2 says ', buf[2]   ! one-sided READ
  end if
end program images

Try it with gfortran -std=f2008 -fcoarray=single images.f90 (syntax and single-image behaviour) and, where the runtime is installed, gfortran -fcoarray=lib images.f90 && cafrun -n 4 ./images (multi-image execution).

Coarray images share no memory but read each other's coarrays

Fig. 1 — Private data per image; square brackets reach across; sync is the rendezvous.

Collectives: reductions across images

Before Fortran 2018, summing a value across images meant explicit per-image reads in a loop, guarded by sync all. The collective subroutines — co_sum, co_min, co_max, co_broadcast and the general co_reduce — fold a value across every image in one call, with the synchronization built in. Every image calls them together; after the call, each image holds the combined result:

program co_sum_demo
  implicit none
  integer :: me, n
  real :: partial
  me = this_image()
  n = num_images()
  partial = real(me * me)          ! each image contributes its own slice
  call co_sum(partial)             ! collective: SUM over all n images
  if (me == 1) then
     print '(a,f8.0)', 'sum of squares across images = ', partial
  end if
end program co_sum_demo

co_broadcast(value, source_image) publishes one image's value to everyone — the standard way to distribute a solution vector or a parameter set at a checkpoint. co_min/co_max reduce error norms and residuals across images, giving every image the global stopping criterion for an iterative solver.

The classic application: distributed arrays

Real coarray code distributes a large array over images — each image owns a contiguous slice, and the [ ] cosubscript reads the neighbours. For a 1-D domain decomposition of a time-dependent simulation, the pattern is: write my slice, sync all, read the two halo values from the neighbouring images' coarrays, compute, repeat. All communication is one-sided and data-flow spelled directly in the source. The MPI lesson contrasts the identical algorithm expressed with explicit message passing — the concepts transfer 1:1, which is precisely why the choice between them is a tooling decision more than a learning decision.

Compiler reality check. gfortran implements coarrays thoroughly (-fcoarray=single for one image; -fcoarray=lib for the multi-image runtime); Cray/Intel have the deepest coarray implementations in HPC production. Teams and the full event model are F2018 and still maturing in free compilers — read the release notes, and when a group feature lags, sync images is the portable fallback.

Synchronization: sync all, sync images, teams

Parallel programs misbehave when an image reads data before the writer published it. sync all is the full barrier — nobody passes until everyone reached it. sync images(list) syncs only with named partners, cheaper on big ensembles. Teams (F2018) partition images into groups that synchronize internally while the whole ensemble proceeds independently — the scalable pattern for million-image runs; events (post/wait) express producer-consumer handoffs that barriers would over-synchronize.