Elixir: Performance & Observability
What Costs What
| Operation | Cost | Practical note |
|---|---|---|
| Spawn a process | ~1–3 µs, ~2.6 KB | spawning per request is normal design |
| Message to another node | TCP hop, copied | design for locality first |
| Append to a list | O(n) | prepend + Enum.reverse instead |
| Map update | O(log n) + structural sharing | fine for millions of keys |
| Pure numeric loop | 10–100× slower than Rust/C | push to Nx/NIFs — see Numerical Elixir |
| Large binary handling | refc heap, ref-counted, shared | slicing binaries is cheap; accreting with <> is not — use IO data |
Shared State That Scales
A GenServer serializes access — at 10k reads/second it becomes the bottleneck. Two escape hatches, both still immutable-data-friendly:
# ETS: an in-memory table outside any process. Readers never block.
table = :ets.new(:cache, [:set, :public, read_concurrency: true, write_concurrency: true])
:ets.insert(table, {:user_42, %{name: "Ada"}})
:ets.lookup(table, :user_42) # microsecond reads, from any process
# :persistent_term — global constants updated rarely, read at memory speed.
:persistent_term.put(:config, %{ttl: 300})
:persistent_term.get(:config)
# Caution: an update triggers a global GC scan — config, not counters.
Profiling
# Benchee — benchmark with warm-up, statistics, and comparisons.
Benchee.run(%{
"map+reverse" => fn list -> Enum.reverse(Enum.map(list, &(&1 * 2))) end,
"comprehension" => fn list -> for i <- list, do: i * 2 end
# Benchee reports ips, avg/dev time, and memory per run.
}, input: %{"10k" => Enum.to_list(1..10_000)})
# :eprof — who calls whom, and how much time each function takes.
:eprof.start()
:eprof.profile(fn -> MyApp.HeavyJob.run() end)
# :erlang.trace / recon — production-safe tracing of specific processes.
:recon_trace.calls({MyApp.HeavyJob, :run, 0}, 10)
Measure in this order: throughput under realistic concurrency (a load script), then queue times (where requests wait), then CPU inside a step. Optimizing a function that is 1% of wall time is the classic waste.
Telemetry and Observability
:telemetry.execute([:my_app, :job, :stop], %{duration: duration}, %{job: "ingest"})
# Attach handlers anywhere (libraries like Phoenix, Ecto, Oban already
# emit standard events — attach, don't reinstrument):
:telemetry.attach("job-metrics", [:my_app, :job, :stop], fn _event, measurements, metadata, _ >
Logger.info("job=#{metadata.job} took=#{measurements.duration}µs")
end, nil)
The ecosystem standard: telemetry events → telemetry_metrics aggregates →
Prometheus/LiveDashboard. Phoenix LiveView apps also ship LiveDashboard out of the
box — live process queues, ETS tables, and metrics with no extra infrastructure.
NIFs and When to Escape
When a hot path stays slow — parsing gigabytes, cryptography, ML inference — native code enters as a NIF (native function call, linked into the VM) or a port (separate OS process). Rules: NIFs must be fast (they block a scheduler thread; dirty schedulers help but are not a license to stall), and Rustler is the standard way to write them safely. The rest of the system never knows — the NIF looks like a normal Elixir function.
Practice
- Benchmark the same transformation with Benchee twice — once on a list, once on a stream — and read the memory column.
- Replace one GenServer-backed cache with an ETS table and measure p99 latency under load.
- Attach a telemetry handler to Ecto's
[:my_app, :repo, :query]events and log queries slower than 100ms.