LLVM IR

LLVM IR is the language-independent intermediate representation at the heart of the LLVM compiler toolchain. It is a low-level, typed, single-assignment language — a DSL written by compilers, for compilers.

Purpose

LLVM IR is the “middle end” of a compiler: front-ends (Clang, rustc, swift, zig, julia) lower source code into IR; a pipeline of optimization passes rewrites it; back-ends lower it to machine code for dozens of architectures.

The Problem It Solves

Writing N compilers for M architectures is N×M work. One well-designed IR reduces that to N front-ends plus M back-ends, and it gives every language the same powerful optimizer. IR exists in three interchangeable forms: readable text, binary bitcode, and in-memory objects — which is why tools such as opt, llc, and lli can treat IR as a manipulable artifact.

Where It Fits

IR is a toolchain DSL, not an application language. It is explicit by design: static single assignment (each value is defined once), typed values, basic blocks, and phi nodes to merge values from different control-flow paths. Written by hand it is tedious; generated by a front-end it is the perfect substrate for analysis and optimization.

History

LLVM began as a doctoral project and became the industry’s default compiler infrastructure.

Origins

Chris Lattner designed LLVM as part of his PhD at the University of Illinois (2002–2004), with the thesis LLVM: A Compilation Framework for Lifelong Program Analysis & Transformation. Version 1.0 shipped in 2003; the novel idea was keeping IR available throughout a program’s lifetime — enabling link-time and runtime (JIT) optimization.

Milestones

  • 2005–2007 — Apple invests in LLVM and Clang; Xcode adopts them.
  • 2011–2015 — Rust and Swift choose LLVM as their backend; Julia adopts it for JIT compilation.
  • 2016+ — the LLVM Foundation formalizes governance; six-month releases continue; MLIR layers on top for even higher-level IR experiments.

Current Status

Extremely mature and still central: Clang, rustc, swiftc, zig, Julia, and many GPU toolchains lower through it. Licensed under Apache 2.0 with the LLVM exception, with backing from Apple, Google, ARM, Qualcomm, and others.

Stage

LLVM IR is battle-tested infrastructure; it changes slowly and compatibly.

Maturity

Extremely mature. The text format is stable for practical purposes, and every release adds back-end and pass improvements without breaking the language contract that front-ends depend on.

Governance & Maintenance

Managed by the LLVM Foundation with corporate backing; six-month release trains; an open review process (Phabricator/GitHub) that has trained a generation of compiler engineers.

Popularity & Usability

As infrastructure, LLVM is everywhere; as a language, IR is only written by compiler developers.

Adoption

Underpins Clang, rustc, swiftc, zig, Julia, and many research and production compilers. Graduate compiler courses and tools like Compiler Explorer (Godbolt) make IR visible to a broad audience.

Learning Curve

Steep but well charted: the LLVM Language Reference is precise, and “LLVM for Grad Students” is a famous quick start. The hard parts are SSA and phi nodes, and thinking in terms of basic blocks and explicit control flow.

Tooling

First-class: clang -S -emit-llvm exports IR from C/C++; opt runs optimization passes; llc lowers to assembly; lli executes IR directly; llvm-dis/llvm-as convert between text and bitcode.

Use Cases

IR is the substrate for anything that compiles, optimizes, or inspects code.

Primary Domains

  • Building new compilers and language front-ends that need a production backend.
  • JIT runtimes — Julia, and many database/GPU runtimes compile IR at runtime for near-native speed.
  • Static analysis and research on optimization passes (students and production teams alike).
  • GPU toolchains: NVPTX and AMDGPU back-ends lower IR for CUDA and ROCm.

Strengths

Language independence, a huge optimization library, proven back-ends, readable text form, and a gentle on-ramp from Compiler Explorer.

Weak Spots

Verbose by design, hostile to handwritten production code, and tied to LLVM’s release cadence; IR details (e.g. poison values, opaque pointers in recent versions) demand constant study.

Performance

IR’s performance story is the optimizer’s story.

Execution Model

IR is lowered through a pipeline of canonical passes (mem2reg, instcombine, GVN, loop unrolling and vectorization, and more) before code generation. SSA form makes many of these passes simple and sound; the same source compiles to efficient code on any supported architecture.

Published Claims

LLVM-generated code is consistently measured within a few percent of GCC on standard benchmarks (SPEC-style), and in JIT settings it is the reason Julia can approach C performance. Claims of “LLVM is faster/slower than GCC” vary by workload and version — treat them as workload-specific.

Example

Handwritten LLVM IR is rare, but reading it is essential. Here is a complete runnable program: two functions, one of them main, so lli can execute it.

The IR

; add.ll — a tiny LLVM IR program
define i32 @add(i32 %a, i32 %b) {
  %sum = add i32 %a, %b        ; typed SSA: every value is defined once
  ret i32 %sum
}

define i32 @main() {
  %r = call i32 @add(i32 40, i32 2)
  ret i32 %r                   ; exit code 42
}

How to Run

With the LLVM tools installed (distro package llvm, or via clang):

lli add.ll            # interpret the IR: runs main, exits with 42
echo $?               # prints 42
llc add.ll -o add.s   # lower IR to x86-64 assembly for study
clang -S -emit-llvm add.c -o add.ll   # see IR generated from C

Notice SSA in action: %sum is assigned exactly once, and every value carries its type (i32). That discipline is what lets the optimizer reason about the program safely.

Learn More

Official sources and free materials; the full categorized catalog is on the References & Downloads page.

Official Docs & Downloads

Learning Material