Design Considerations

Before a lexer, before a grammar: decide whether inventing a language is justified, which ideas are reusable, and what the language will cost to design, build, and maintain forever. This chapter is the decision layer above syntax, grammar, and parsing.

The Decision Funnel

How to Read the Funnel

The whole design problem fits in one picture: start from a measured problem, decide reuse or invention, then choose the paradigms, semantics, and syntax you will deliberately keep or change.

Decision funnel: problem, reuse-or-invent, paradigms and semantics, syntax and grammar, implementation, tooling and AI assistance, then measure

Figure 1 — The design decision funnel: reuse first, invent only with justification, disrupt only where the cost is paid back.

Why Design a Language?

Languages are long-lived capital. Building one is justified only when the motivation is real and the cost can be amortized.

Motivations

  • Abstraction gap: the domain has vocabulary and rules no general language expresses directly — SQL for data, Verilog for hardware, AutoLISP for CAD.
  • Productivity: remove boilerplate and encode invariants so the same task takes less, safer code.
  • Control: add safety rules a host language cannot enforce — Rust-style ownership, Kotlin null-safety.
  • Automation and interop: let users script your tool (Emacs Lisp, AutoLISP) or describe what it consumes.
  • Learning and research: language design teaches computation; many great systems started as experiments.

Hype vs. Real Need

Suspect every motivation that is not a problem with a name. “Nicer syntax”, “more modern”, or “I could write a language” describe style, not need. A defensible reason names the users, the workflow, and the cost the language removes.

Amortizing the Investment

Design, implementation, tooling, documentation, and maintenance are a permanent cost. The language pays for itself across users, across years, or as the product surface of a tool. If none of those hold, an embedded or configuration layer probably serves better.

When to Invent & What Is Reusable

Inventing is the default instinct and often the wrong one. Use the signals and the toolkit below to decide deliberately.

Signals to Invent

  • No existing language, DSL, or format fits without fighting it.
  • The domain needs safety rules no host language provides (borrowing, nullability, temporal rules).
  • The investment amortizes across many users or one long-lived product.
  • The language is part of your product’s surface — customers script against it.

Reuse Before Inventing

  • Configuration formats: JSON, YAML, and TOML cover most data-shaped needs.
  • Embedded DSLs: macros, builders, fluent APIs, and decorators often look like a language already.
  • Host extension layers: annotations, templates, and plugin interfaces add vocabulary without a parser.
  • Interoperability: an FFI or a substrate (JVM, .NET, WASM) gives you an ecosystem for free.

The Reusable Toolkit

You never start from nothing: paradigms (imperative, functional, object-oriented, logic, declarative); conventions (expression syntax, operator precedence, naming); rules (scoping, typing, evaluation order, memory); and machinery (interpreters, stack machines, bytecode, LLVM IR — the later phases of this roadmap). Count what you reused; if the count is small, ask why you are writing a parser at all.

Paradigms, Conventions & Rules

Choosing a paradigm is choosing a set of reusable ideas; conventions and rules decide how familiar or how foreign the language feels.

Paradigm Inventory

  • Imperative: statements change state; the mental model maps to hardware (C, Forth).
  • Functional: values and pure transformations; concurrency and reasoning benefit (Lisp, Clojure, Haskell).
  • Object-oriented: messages and encapsulation organize large systems (Java, C#).
  • Logic/declarative: you state what, the engine searches how (Prolog, Datalog, SQL, MiniZinc — Phase 6).
  • Dataflow/actor: streaming and message passing fit reactive and concurrent domains.

Conventions Users Expect

Familiar syntax shapes, C-family roots, conventional operators, and readable keywords lower the learning curve. Conventions are also what AI models have seen billions of times: a conventional language gets better generated code and better tooling. Deviating from a convention is a feature only when it buys a rule the domain needs.

Semantic Rules to Decide

Write these down before any code: lexical or dynamic scoping; static, dynamic, or gradual typing; strict or lazy evaluation; reference, value, or ownership memory semantics; exceptions or result values. Each choice constrains every later chapter of this roadmap.

Disruption vs. Compliance

Every design decision is a bet between staying familiar and changing the rules. The bet pays off only when the disruption fixes a real bug class or a real cost.

Why Disrupt

Rust’s ownership removes memory bugs at compile time; Kotlin’s non-null-by-default removes null-pointer crashes; Swift’s optionals make absence explicit. Each is a rule change that eliminates an entire class of defects — the strongest argument a language can make.

The Cost of Disruption

  • Tooling, libraries, and community must be rebuilt or re-taught.
  • Migration converts working systems and trains every developer.
  • Adoption slows because the language looks “wrong” to veterans.
  • A failed disruption wastes the whole amortization case.

When to Comply

Compliance wins adoption, interop, and AI-tooling quality. It is a decision, not a failure of imagination: keep the familiar 90% of conventions, and spend your disruption budget on the few rules that matter for your domain.

Critical Thinking

Language design filters opinions through evidence. These checks keep the project honest.

Decision Checklist

  1. Name the problem and how you will measure success.
  2. Check reuse first: format, embedded DSL, existing language.
  3. Design the smallest language that solves it.
  4. Write example programs before implementing anything.
  5. Declare which paradigms, conventions, and rules you reuse.
  6. Decide disruption vs. compliance explicitly, rule by rule.
  7. Plan the exit: can users migrate away if the language fails?

Red Flags

Suspicious starts: syntax-first obsession (“just nicer syntax”); solving a community or process problem with a language change; claims without a benchmark (“blazingly fast”); copying a language with no delta; and design docs that never mention users.

Problems to Resolve Early

During the design phase, settle these four problem families or they will settle you during implementation.

Syntax and Grammar

Unambiguous grammar, clear precedence, and error recovery are the Phase 1 material (Syntax, Grammar, Parsing). A grammar that fights you is a design signal, not a tooling problem.

Semantics

Scoping, typing, evaluation order, and memory rules must be written down before code generation; most failed languages die from undefined semantics.

Errors and Tooling

Diagnostics, REPL, formatter, and debugger decide whether beginners survive. A language without a way to ask “what is happening?” is unteachable.

Ecosystem and Docs

Packaging, host interop, and documentation with commented examples are part of the language contract — not an afterthought.

The AI Landscape

Who writes the code is changing, and that changes language design.

AI-Assisted Code Generation

Models produce the most correct code in conventional, unambiguous languages with abundant training examples. A precise grammar, commented docs, and plenty of examples measurably improve generated output — treat machine-readable specs (EBNF, schemas) as part of the design.

Verification and Comments

AI writes plausible code; static checks and tests catch model errors early. The commented-example rule of this roadmap is truth in advertising for humans and models alike.

AI Domain Languages

Phase 4 of this roadmap (Triton, Mojo) exists because AI workloads need domain languages. New DSLs for prompts, evals, and model definitions are arriving; design them with the same funnel: reuse, justify, comply where possible.

Worked Mini-Example

A team automates room bookings. Before inventing, the funnel asks: is there a problem with a name, and does reuse fail?

The Requirements

Each booking names a room, a time window, and features (capacity, projector). Rules: no double booking; suggest alternatives on conflict.

Decisions Made

  • Reuse won: the language is a small embedded DSL parsed by a 20-line reader, not a compiler project.
  • Compliance: English keywords and YAML-like lines — nothing exotic.
  • Deliberately small: no loops, no expressions; expressiveness is a choice, not a gap.

The DSL (commented)

# booking.req — one request, deliberately small vocabulary
room 3A                 # the room identifier (a name, not an expression)
time 09:00-10:30        # fixed window; no arithmetic on purpose
capacity 6              # seats required (an integer payload)
projector yes           # feature flag
on_conflict suggest     # policy word chosen from a short menu

Grammar Sketch (commented EBNF)

(* booking.req — every line is keyword + payload; no nesting *)
request      = room_line, time_line, { feature_line }, on_conflict ;
room_line    = "room", name ;
time_line    = "time", clock, "-", clock ;
feature_line = "capacity", number | "projector", flag ;
on_conflict  = "on_conflict", ( "suggest" | "fail" ) ;

Tiny Reader (commented Python)

# reader.py — the whole interpreter for the booking DSL
def parse(lines):
    # Return a dict of keyword -> payload for each line.
    data = {}
    for line in lines:
        line = line.split("#")[0].strip()   # drop comments, trim spaces
        if not line:
            continue                        # skip blank lines
        key, _, rest = line.partition(" ")  # first token is the keyword
        data[key] = rest or True
    return data                             # caller validates the values

In a few lines the team shipped the DSL; a full compiler would have failed the amortization test. This is the funnel working. Python keeps every example readable no matter which language you end up using; when the DSL grows into a real compiler, the DSL Implementation Languages phase covers the production paths: ANTLR for grammar-driven parsing, Racket for language design, and OCaml for a type-safe core.

Next Steps

Continue Phase 1

With the decision made, the mechanics follow: Syntax & Notation, Grammar & Rules, and Parsing & Trees turn the design into a specification, and Phase 2 turns the specification into an implementation.

Resources