Elixir: Supervision Trees

A supervisor is a process whose only job is watching other processes and restarting them when they die. Supervisors supervising supervisors form a tree — and that tree is the fault-tolerance architecture of the whole system, visible in one place: the application module.

The First Supervisor

defmodule MyApp.Supervisor do
  use Supervisor    # the behaviour for static trees

  def start_link(init_arg), do:
    Supervisor.start_link(__MODULE__, init_arg, name: __MODULE__)

  @impl true
  def init(_init_arg) do
    # children: a list of child specs. ORDER = boot order, reversed = shutdown.
    children = [
      # Start order: registry first, then cache, then workers.
      {Registry, keys: :unique, name: MyApp.Registry},
      {MyApp.Cache, name: MyApp.Cache},
      {Task.Supervisor, name: MyApp.TaskSup}
    ]

    # one_for_one: a crashed child restarts ALONE, the rest keep running.
    Supervisor.init(children, strategy: :one_for_one)
  end
end

# Manual boot for scripts and experiments (real apps start it via the
# application callback — see the next lesson):
Supervisor.start_link(MyApp.Supervisor, [])

Restart Strategies

A supervision tree: the application supervisor watches a Registry, a Cache and a worker supervisor, which watches three workers with one_for_one
Fig. 1 — Restart cascades stop at supervisor boundaries: each supervisor restarts only what its strategy dictates.
StrategyOn a child crashUse when
:one_for_oneonly that child restartsindependent children (the default choice)
:one_for_allall children restartchildren are useless without each other
:rest_for_onethe crashed child and everything started after itlater children depend on earlier ones

Restart intensity protects the system from crash loops: max_restarts failures within max_seconds (defaults: 3 in 5) take the supervisor itself down — escalating the failure one level up the tree. That "give up and let the parent decide" rule is what makes transient bugs self-healing instead of wedge-shaped.

DynamicSupervisor

Static trees know their children at compile time. When children appear at runtime — one process per user connection, per game, per upload — DynamicSupervisor starts and supervises them on demand, and pairs with a Registry for name-based lookup.

# In the tree: {DynamicSupervisor, name: MyApp.SessionSup}

# Later, at runtime — "supervise a session for user 42":
{:ok, pid} =
  DynamicSupervisor.start_child(MyApp.SessionSup, %{
    id: {MyApp.Session, 42},
    start: {MyApp.Session, :start_link, [[user_id: 42]]},
    restart: :transient          # restart on abnormal exit only
  })

DynamicSupervisor.count_children(MyApp.SessionSup)
DynamicSupervisor.terminate_child(MyApp.SessionSup, pid)

restart: :temporary and :transient matter here: a crashed session should restart, but a finished session should not resurrect itself.

Designing Fault Domains

The tree is your failure architecture. Rules that hold up in production:

  • One concern per supervisor. A crashed HTTP pool should not restart your DNS resolver.
  • Restart cheap things, rebuild expensive things. A cache server can be :permanent; a connection pool with a slow warm-up wants a strategy that avoids flapping.
  • State that must survive a crash lives outside the process — ETS table owned by a separate "guardian" process, or persisted with Ecto.
  • Watch your own intensity limits: defaults are sensible; tune max_restarts/max_seconds only with a reason.

Practice

  1. Start the supervisor above in iex, find a child pid with Supervisor.which_children/1, and kill it — watch the restart.
  2. Change the strategy to :one_for_all and repeat; explain the difference in what restarted.
  3. Build a DynamicSupervisor that hosts one MyApp.Session GenServer per user id, then terminate one session by name.

Next: OTP Applications & Releases