Triton

Triton is a Python-based language and compiler for writing GPU kernels productively. You program in blocks (tiles), not threads, and Triton generates the low-level GPU code — delivering CUDA-class performance without hand-writing CUDA.

Purpose

Triton is an AI-domain DSL: it exists to make custom deep-learning kernels (attention, fused matmuls, activations) fast to write and fast to run on modern GPUs.

The Problem It Solves

Hand-written CUDA gives full performance but is slow to write, hard to port, and requires experts. Using only prebuilt operations (cuBLAS, cuDNN, PyTorch ops) means fusing and customizing is impossible. Triton offers a middle path: a Python kernel language whose compiler handles thread scheduling, memory coalescing, and vectorization for you.

Where It Fits

Triton sits between high-level frameworks and raw CUDA. Each kernel declares a grid of program instances, each operating on a block of data with explicit loads and stores; the compiler maps blocks onto GPU hardware. Because it compiles through MLIR to LLVM, the same source targets NVIDIA and AMD GPUs.

History

Triton has the rare trajectory of an academic project adopted as production infrastructure.

Origins

Philippe Tillet built Triton during doctoral work at Harvard, publishing the design in the PLDI 2019 paper Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations.

Milestones

  • 2021 — OpenAI adopts and open-sources Triton for internal GPU kernel development.
  • 2022–2023 — PyTorch 2.0’s torch.compile (Inductor) emits Triton kernels by default on NVIDIA GPUs, making Triton the default custom-kernel path for the PyTorch ecosystem.
  • 2024+ — the community triton-lang project extends backend support (AMD, experimental CPU) while the OpenAI repository remains the canonical fast-moving implementation.

Current Status

Actively developed and production-critical: whenever you run an optimized PyTorch model, Triton kernels are likely inside it. MIT-licensed, with tutorials, a Python API, and documentation at triton-lang.org.

Stage

Triton is young but already load-bearing: PyTorch’s default compiled path runs through it.

Maturity

Fast-moving pre-1.0-style churn with stable core semantics. The Python API (triton.language) and the MLIR-based compiler change quickly; production users pin versions.

Governance & Maintenance

Primarily maintained by OpenAI, with a growing community (triton-lang) driving additional back-ends (ROCm/AMD, experimental CPU). MIT-licensed and open to contributions.

Popularity & Usability

Within machine learning engineering, Triton is close to a standard skill.

Adoption

torch.compile emits Triton by default on NVIDIA GPUs, so every PyTorch 2.x user runs Triton-generated kernels. Kernel libraries for attention, quantization, and sampling are increasingly written in Triton across the ecosystem.

Learning Curve

Moderate if you know tensors and PyTorch: no thread IDs or block indexing to manage, just tl.program_id, tl.arange, and tile-wide loads/stores with masks. The mental model (blocks parallelize over the data) is the main new idea.

Tooling

pip install triton, a @triton.jit decorator, autotuning helpers (triton.autotune), plus profiling against the native profilers. Debugging is less mature than CUDA’s — expect to reason about generated code via print or disassembly.

Use Cases

Anything that needs a custom, fast GPU kernel without hand-written CUDA.

Primary Domains

  • Fused attention (FlashAttention-style) and GEMM/MLP fusions.
  • Custom activations, normalization, RoPE, and quantization kernels.
  • Everything torch.compile cannot fuse, or that needs a specialized layout.

Strengths

Python productivity, automatic coalescing and vectorization, portable source across GPU vendors, and first-class integration with PyTorch tensors.

Weak Spots

Younger debugging story, rapid version churn, and gaps for exotic hardware features; if you need bleeding-edge GPU instructions, raw CUDA remains the fallback.

Performance

Triton’s entire reason for existing is closing the gap to hand-tuned CUDA.

Execution Model

A Triton kernel runs a grid of program instances; each instance works on a tile defined by tl.arange offsets. The compiler chooses how to split registers, vectorize memory access, and schedule warps — the jobs that make CUDA kernels hard to write.

Published Claims

The PLDI 2019 paper reported Triton kernels matching or beating hand-written CUDA on several tiled neural-network workloads (attention and matmul style). In practice, Inductor uses Triton to recover most of cuBLAS/cuDNN-level performance with Python-level effort; the honest benchmark rule is: measure per kernel, because tile sizes and masks decide the outcome.

Example

A complete Triton kernel and its host launcher: element-wise vector addition with a mask for the tail of the array.

The Kernel

# add_kernel.py — Triton GPU kernel for element-wise addition
import torch
import triton
import triton.language as tl

@triton.jit
def add_kernel(x_ptr, y_ptr, z_ptr, n, BLOCK: tl.constexpr):
    # Each program instance handles one BLOCK-sized tile.
    pid = tl.program_id(0)                 # which tile am I?
    offs = pid * BLOCK + tl.arange(0, BLOCK)
    mask = offs < n                        # guard elements past the end
    x = tl.load(x_ptr + offs, mask=mask)   # masked load: out-of-range -> 0
    y = tl.load(y_ptr + offs, mask=mask)
    tl.store(z_ptr + offs, x + y, mask=mask)

# Host side: allocate tensors on the GPU and launch a grid of tiles.
x = torch.randn(1024, device="cuda")
y = torch.randn(1024, device="cuda")
z = torch.empty_like(x)
grid = (triton.cdiv(1024, 256),)          # 4 tiles of 256 elements
add_kernel[grid](x, y, z, 1024, BLOCK=256)
torch.testing.assert_close(z, x + y)      # compare with PyTorch

How to Run

python -m pip install triton torch   # install both
python add_kernel.py                 # runs the assertion on an NVIDIA GPU

The mask is the key idiom: tensors rarely divide evenly into tiles, so every load/store carries a predicate. The compiler turns this tile program into coalesced, vectorized GPU instructions — no explicit threads anywhere.

Learn More

Official sources and free materials; the full categorized catalog is on the References & Downloads page.

Official Docs & Downloads

Learning Material