Stan

Stan is a probabilistic programming language for Bayesian statistical modeling: you declare a model (data, parameters, priors, likelihood), and Stan’s Hamiltonian Monte Carlo engine draws posterior samples — from any interface: R, Python, Julia, or the shell.

Purpose

Stan lets statisticians write the model, not the algorithm. Its compiler turns a declarative model into an efficient C++ sampler, so Bayesian inference is expressible in a page of code.

The Problem It Solves

Fitting hierarchical and non-standard models by hand (deriving full conditionals, writing samplers) is error-prone and slow. Stan separates model specification from inference: you declare blocks for data, parameters, and the joint model; the engine computes gradients and samples with Hamiltonian Monte Carlo (HMC), made practical by the no-U-turn sampler (NUTS). Diagnostics (R-hat, effective sample size) come built in.

Where It Fits

Stan is this phase’s Bayesian DSL: it complements R and Python (statistics/data), Julia (numeric speed), and Wolfram (symbolic) by owning posterior inference. It is the modern successor to BUGS/JAGS in the same modeling niche.

History

Stan was built by statisticians who believed HMC could replace hand-written Gibbs samplers.

Origins

The Stan Development Team, guided by Andrew Gelman and led in implementation by Bob Carpenter and colleagues, released Stan 1.0 in 2012. It grew out of Columbia’s statistics community and the C++ numerical tradition.

Milestones

  • 2012–2016 — rapid adoption in academia; interfaces (RStan, PyStan, CmdStan) mature; NUTS becomes the default sampler.
  • 2017+ — CmdStanR/CmdStanPy modernize the user experience; case studies and the prior-choice wiki formalize best practice.
  • Today — active 2.x releases, an enormous forum/mailing-list knowledge base, and the reference implementation for probabilistic-programming performance comparisons.

Current Status

Mature, BSD-3-licensed, and the de-facto standard for Bayesian modeling in research and industry (epidemiology, pharma, A/B testing, psychometrics). Interfaces for R, Python, Julia, and the command line all share one model language.

Stage

Stan is mature and in heavy production use across statistics and applied science.

Maturity

Fully mature: a stable model language, hardened C++ runtime, and a decade of diagnostics tooling. Versioning stays in the 2.x line, with additive changes and careful compatibility.

Governance & Maintenance

The Stan Development Team (open to contributors) maintains the compiler and interfaces; development is public and BSD-3-licensed. The Stan forums provide long-lived community support.

Popularity & Usability

If Bayesian modeling has a lingua franca, it is Stan.

Adoption

Widely used in epidemiology, pharma, psychometrics, economics, and Bayesian machine learning; standard in graduate statistics curricula; the reference point for HMC-based inference tools.

Learning Curve

The model DSL itself is small: data, parameters, model blocks. The real learning curve is Bayesian statistics: priors, posterior diagnostics (R-hat, ESS, divergences), and reparameterization. Interface choice (CmdStanR, PyStan, cmdstanpy) is well documented.

Tooling

CmdStan (CLI), RStan, PyStan/CmdStanPy, Stan.jl (Julia), plus the broader ecosystem: bayesplot, rstanarm (pre-built models), brms (formula interface), and bridgesampling.

Use Cases

Stan is the right tool when the uncertainty in the answer matters.

Primary Domains

  • Hierarchical (multilevel) models: students within schools, patients within hospitals, countries within regions.
  • Meta-analysis and evidence synthesis in medicine and social science.
  • Bayesian A/B testing and marketing measurement.
  • Epidemiological forecasting, psychometrics (IRT), and state-space time series.

Strengths

Declarative models, HMC/NUTS efficiency, excellent diagnostics, and one model language across R/Python/Julia/CLI.

Weak Spots

Both programming: discrete parameters need marginalization tricks; very large datasets slow HMC; and the DSL deliberately excludes imperative constructs, which surprises programmers.

Performance

Stan’s performance metric is posterior samples per second.

Execution Model

The model is compiled to C++ with automatic differentiation; NUTS adaptively chooses step sizes to explore the posterior efficiently, so the sampler spends effort only where probability mass lives.

Published Claims

The Stan community’s published case studies and comparisons with JAGS/BUGS consistently show HMC needing far fewer, higher-quality samples — the “effective sample size per second” win. The honest caveats: performance is model-shaped (posteriors with strong correlations need reparameterization), and pathfinder/optimized sub-samplers exist for hard cases.

Example

The canonical “coin flip” model: infer the probability of heads from observed outcomes, then fit it from R through CmdStanR.

The Stan Model

// bernoulli.stan — infer a success probability from coin flips
data {
  int<lower=0> N;                    // number of trials
  array[N] int<lower=0, upper=1> y;  // observed outcomes (0/1)
}
parameters {
  real<lower=0, upper=1> theta;      // success probability
}
model {
  theta ~ beta(1, 1);           // uniform prior on (0,1)
  y ~ bernoulli(theta);         // likelihood
}

How to Run (CmdStanR)

# fit_bernoulli.r — fit the model from R
library(cmdstanr)
mod <- cmdstan_model("bernoulli.stan")
fit <- mod$sample(data = list(N = 8, y = c(1,0,1,1,1,0,1,0)),
                  seed = 123, chains = 4)
fit$summary("theta")            # posterior mean and credible intervals

The data/parameters/model block structure is the whole language in miniature: no loops for sampling, no solver choice — you declare the joint distribution and Stan computes the posterior.

Learn More

Official sources and free materials; the full categorized catalog is on the References & Downloads page.

Official Docs & Downloads

Learning Material