Code Generation

Native codegen is the deepest rung of the IR ladder: instructions, registers, calling conventions, and object code. Most DSLs never climb it — they hand the IR to LLVM or a host compiler instead — but knowing the rung explains every back-end decision.

Instruction Selection

Map each IR operation to target instructions. Choice for t1 = t2 * 3 on x86-64: imul t1, t2, 3 (fold the constant into the instruction) — the same job the Bytecode emitter did, now with target details like addressing modes.

Coverage, Not Beauty

A correct-but-slow selector beats a clever one that misses cases; compilers generate obscure code far more often than they make it beautiful.

Register Allocation

Registers are fast and few; memory is slow and large. Allocation decides which IR values live where, and what spills (moves to stack) when a function has more live values than registers.

One Practical Algorithm

Linear scan assigns registers in program order, expiring values that are no longer live (liveness from liveness analysis) and spilling when a register is unavailable. Simple, fast, perfectly adequate for DSL-sized functions.

Calling Conventions

ABI Basics

An ABI contract: which registers pass arguments, who saves callee-saved registers, and how the stack is aligned and unwound. If you emit only your own calls, a minimal home-grown convention suffices; the moment you call the host or the OS, you must obey theirs.

When to Go Native (And When Not)

The LLVM Escape Hatch

Native codegen pays off when there is a hot loop to speed up and no foreign function interface will do. Otherwise the correct engineering answer is: emit LLVM IR (Phase 3) and let a battle-tested back-end own registers, ABIs, and peepholes. Writing your eighth register allocator is rarely the DSL’s best investment.

Example: IR to Instructions

One selector rule and one comment shows the entire native rung.

select.py (commented)

# select.py -- map the three-address IR onto a target
#   IR:        t1 = t2 * t3          t4 = t1 + 1
#   x86-64:    mov  rax, [t2]        add  rax, t1[+1 is folded]
#              imul rax, [t3]
#              mov  [t1], rax

How to Read It

Selection, then registers, then the ABI — each step adds target detail. If this looks like a lot of system code, that is correct: it is exactly the machinery LLVM owns (Phase 3) so your DSL need not.

Next Steps

Continue Phase 2

Speed has one more chapter: Code Optimization, then Phase 3 (LLVM IR, Triton, Mojo) takes these ideas to production compilers.

Resources