Syntax & Notation

Syntax is the surface form of a language: how characters group into tokens, how tokens group into phrases, and the notational conventions readers expect. A DSL’s syntax is its first impression — and its cheapest place to get design wrong.

Lexical Structure

Before structure, the characters themselves must be divided into meaningful pieces.

Lexical pipeline: source text becomes a stream of tokens through the lexer

Figure 1 — The lexer turns characters into tokens; comments and whitespace are consumed here.

Tokens and Lexemes

A lexeme is the text (“3A”); a token is its kind plus value (NAME(3A)). DSL token kinds are usually few: keywords, names, numbers, strings, punctuation. Deciding them early fixes the vocabulary users will type.

Whitespace and Comments

Choose deliberately: insignificant whitespace (most languages) or significant (Python, Make, YAML), and accept only one comment syntax (#, //, or --; never several). Comments are promise to maintainers; lexers that skip them keep the grammar clean.

Names and Literals

Decide identifier rules (letters, case, underscores, Unicode), number literals (integer vs decimal, hex/binary prefixes), and string quoting. Conventions reduce surprise: if the ecosystem around you uses snake_case and double quotes, breaking that convention buys nothing.

Notation Conventions

Notation is the idiom of expressions: which operator stands where, and how tightly it binds.

Expression Forms

  • Infix: a + b — the familiar default; requires precedence rules (»Grammar & Rules»).
  • Prefix: (+ a b) — Lisp-style; uniform and unambiguous, cheap to parse.
  • Postfix: a b + — stack style (Forth); powerful, hostile to beginners.

DSLs rarely need all three. Pick one primary form and keep the rest out of the vocabulary.

Precedence and Associativity

Infix syntax forces two conventions: operator precedence (which operator binds first) and associativity (left or right if operators repeat). Neither can be left implicit: users assume the arithmetic defaults, so reuse them unless the domain says otherwise.

Surface vs. Structure

Syntax is not meaning. The concrete form is what the user types; the structure (parse tree) is what the machine builds after the grammar applies. “Syntactic sugar” is any surface form that maps onto an existing structure — sugar is fine as long as its grammar stays unambiguous.

Readability Rules

Keywords read better at phrase starts (room 3A, not 3A room); a line should fit on one screen; and the most common action should have the fewest keystrokes. Every rule above readability costs grammar complexity — budget it.

Consistency Wins

One comment marker, one string quote, one number literal spelling. Users forgive unusual notation faster than inconsistent notation.

Design Hints

Keep the Lexer Small

If the lexer needs context to classify a token (is a/b a path or a division?), the design is leaking complexity into Phase 1. Prefer a token kind that is a stable, context-free word.

Tokens Match Errors

Diagnostic messages reference tokens, so token boundaries must align with what users type. A token that can merge two meanings (“and both a keyword and a name”) makes every error message ambiguous.

Example: A 25-Line Lexer

A real DSL begins with a real lexer. This one tokenizes the booking line room 3A — enough to show every idea from this chapter.

The Lexer (commented Python)

# lexer.py — turn "room 3A # note" into tokens
import re

# one rule per token kind; order decides priority (keywords first)
RULES = [
    ("KEYWORD", r"room|time|capacity|projector"),
    ("NUMBER",  r"\d+"),
    ("NAME",    r"[A-Za-z0-9_-]+"),   # identifiers and codes
    ("NEWLINE", r"\n"),
]

def lex(text):
    # Slice TEXT into (kind, value, offset) tokens; skip spaces and comments.
    tokens, pos = [], 0
    while pos < len(text):
        if text[pos] in " #":          # whitespace and comments are noise
            pos += 1
            continue
        for kind, pattern in RULES:
            m = re.match(pattern, text[pos:])
            if m:                      # first rule that matches wins
                tokens.append((kind, m.group(), pos))
                pos += m.end()
                break
        else:
            raise SyntaxError(f"unexpected char {text[pos]!r} at {pos}")
    return tokens

print(lex("room 3A"))   # [('KEYWORD','room',0), ('NAME','3A',5)]

How to Read It

The lexer has exactly the tokens this DSL’s grammar needs, in a fixed priority order; spaces and # comments vanish in one line. If a later chapter needs 09:00, a TIME rule with a better pattern replaces the NUMBER match without touching the grammar.

Next Steps

Continue Phase 1

Syntax defines the words; Grammar & Rules defines how they combine, and Parsing & Trees turns both into a tree the machine can walk.

Resources