Elixir: Performance & Observability

BEAM performance is unusual: single-process numeric work is slow, but system throughput under thousands of concurrent tasks is excellent — if you know where the bottlenecks are. This lesson covers the runtime costs, the escape hatches, and the tools that tell you which one you need.

What Costs What

OperationCostPractical note
Spawn a process~1–3 µs, ~2.6 KBspawning per request is normal design
Message to another nodeTCP hop, copieddesign for locality first
Append to a listO(n)prepend + Enum.reverse instead
Map updateO(log n) + structural sharingfine for millions of keys
Pure numeric loop10–100× slower than Rust/Cpush to Nx/NIFs — see Numerical Elixir
Large binary handlingrefc heap, ref-counted, sharedslicing binaries is cheap; accreting with <> is not — use IO data

Shared State That Scales

A GenServer serializes access — at 10k reads/second it becomes the bottleneck. Two escape hatches, both still immutable-data-friendly:

# ETS: an in-memory table outside any process. Readers never block.
table = :ets.new(:cache, [:set, :public, read_concurrency: true, write_concurrency: true])
:ets.insert(table, {:user_42, %{name: "Ada"}})
:ets.lookup(table, :user_42)          # microsecond reads, from any process

# :persistent_term — global constants updated rarely, read at memory speed.
:persistent_term.put(:config, %{ttl: 300})
:persistent_term.get(:config)
# Caution: an update triggers a global GC scan — config, not counters.

Profiling

# Benchee — benchmark with warm-up, statistics, and comparisons.
Benchee.run(%{
  "map+reverse" => fn list -> Enum.reverse(Enum.map(list, &(&1 * 2))) end,
  "comprehension" => fn list -> for i <- list, do: i * 2 end
  # Benchee reports ips, avg/dev time, and memory per run.
}, input: %{"10k" => Enum.to_list(1..10_000)})

# :eprof — who calls whom, and how much time each function takes.
:eprof.start()
:eprof.profile(fn -> MyApp.HeavyJob.run() end)

# :erlang.trace / recon — production-safe tracing of specific processes.
:recon_trace.calls({MyApp.HeavyJob, :run, 0}, 10)

Measure in this order: throughput under realistic concurrency (a load script), then queue times (where requests wait), then CPU inside a step. Optimizing a function that is 1% of wall time is the classic waste.

Telemetry and Observability

:telemetry.execute([:my_app, :job, :stop], %{duration: duration}, %{job: "ingest"})

# Attach handlers anywhere (libraries like Phoenix, Ecto, Oban already
# emit standard events — attach, don't reinstrument):
:telemetry.attach("job-metrics", [:my_app, :job, :stop], fn _event, measurements, metadata, _ >
  Logger.info("job=#{metadata.job} took=#{measurements.duration}µs")
end, nil)

The ecosystem standard: telemetry events → telemetry_metrics aggregates → Prometheus/LiveDashboard. Phoenix LiveView apps also ship LiveDashboard out of the box — live process queues, ETS tables, and metrics with no extra infrastructure.

NIFs and When to Escape

When a hot path stays slow — parsing gigabytes, cryptography, ML inference — native code enters as a NIF (native function call, linked into the VM) or a port (separate OS process). Rules: NIFs must be fast (they block a scheduler thread; dirty schedulers help but are not a license to stall), and Rustler is the standard way to write them safely. The rest of the system never knows — the NIF looks like a normal Elixir function.

Practice

  1. Benchmark the same transformation with Benchee twice — once on a list, once on a stream — and read the memory column.
  2. Replace one GenServer-backed cache with an ETS table and measure p99 latency under load.
  3. Attach a telemetry handler to Ecto's [:my_app, :repo, :query] events and log queries slower than 100ms.

Next: Numerical & Embedded Elixir