Study Projects

Three programs that pull the whole track together, written with the standard library only. Type each one yourself — do not copy and paste — then break it deliberately and repair it. These are the kinds of programs that teach more than any amount of reading: a text analysis with dictionaries and sorting, a scheduler that uses abstract types and multiple dispatch, and a numerical simulation that implements its own integrator.

Each project is complete as shown and takes about an hour to write and understand. The challenge at the end of each is deliberately open: it is the version of the project that you would meet in real work, where the requirements arrive as a paragraph rather than as a specification.

Project 1 — Word Frequency Counter

Every programmer writes this program once. It is the smallest project that exercises files, strings, dictionaries, sorting, and formatting together — and the smallest one where the obvious implementation is also the slow one if you let it re-read the file.

The design: read the file once, normalise each word, count with a Dict, then sort the entries. Nothing here needs a package, and every step maps onto a lesson from earlier in the track.

# word_frequency.jl — count words in a text file
module WordFrequency

export count_words, top_words, report

"""
    count_words(path) -> Dict{String,Int}

Read the file at `path` once and return word => count.
A word is a run of letters, digits, apostrophes or hyphens; everything
else is a separator. Matching is case-insensitive.
"""
function count_words(path::AbstractString)
    isfile(path) || throw(ArgumentError("no such file: $path"))

    counts = Dict{String,Int}()
    for line in eachline(path)                    # streams: no whole-file copy
        for raw in split(line)
            word = normalise(raw)
            isempty(word) && continue              # punctuation-only token
            counts[word] = get(counts, word, 0) + 1
        end
    end
    return counts
end

"""
Lowercase a token and strip leading/trailing punctuation.
Everything inside the token is kept, so "state-of-the-art" stays one word.
"""
function normalise(token::AbstractString)
    # keep letters, digits, apostrophes and hyphens; drop the rest
    cleaned = filter(c -> isletter(c) || isdigit(c) || c in ('\'', '-'), lowercase(token))
    return strip(cleaned, ['\'', '-'])            # "word." -> "word", "'word'" -> "word"
end

"""
    top_words(counts, n) -> Vector{Pair{String,Int}}

The `n` most frequent words, most frequent first. Ties break alphabetically
so the output is stable between runs — important for tests and for diffs.
"""
function top_words(counts::AbstractDict, n::Integer = 10)
    entries = collect(counts)
    sort!(entries, by = e -> (-e.second, e.first))
    return entries[1:min(n, length(entries))]
end

"""Render a report with counts and a simple bar chart."""
function report(path::AbstractString; top::Integer = 10)
    counts = count_words(path)
    words  = top_words(counts, top)
    total  = sum(values(counts))
    unique_words = length(counts)

    widest = isempty(words) ? 0 : maximum(length(first(w)) for w in words)
    println("file            : ", path)
    println("tokens          : ", total)
    println("distinct words  : ", unique_words)
    println("type/token ratio: ", round(unique_words / max(total, 1), digits = 3))
    println()
    for (word, count) in words
        bar = repeat("#", min(count, 40))         # cap the bar at terminal width
        println(rpad(word, widest), "  ", lpad(count, 6), "  ", bar)
    end
    return counts
end

end # module

# ---------------------------------------------------------------
if abspath(PROGRAM_FILE) == @__FILE__              # only when run as a script
    using .WordFrequency
    isempty(ARGS) && (println(stderr, "usage: julia word_frequency.jl FILE [N]"); exit(2))
    n = length(ARGS) >= 2 ? parse(Int, ARGS[2]) : 10
    WordFrequency.report(ARGS[1]; top = n)
end

Three details carry the lesson. eachline(path) streams the file instead of loading it, so the program survives a 500 MB log. get(counts, word, 0) + 1 is the idiomatic counter — a single hash lookup, no branch. And sorting by a tuple (-count, word) gives the stable ordering that makes the output testable.

Challenge

Extend the program in three steps, in this order. First, add a --min-length option that ignores words shorter than a given length — the option parser belongs in the script block, not in the module. Second, make it read standard input when the path is -, so it can be used in a pipeline: cat book.txt | julia word_frequency.jl -. Third, add a @testset that checks normalise against a table of awkward cases ("End.", "'quoted'", "co-op", "123") and checks that a tie in counts produces alphabetical order. When the third step is easy, the design is right.

Project 2 — Priority Task Scheduler

The second project is about types. A scheduler holds heterogeneous work items, orders them by priority and deadline, and runs them. The interesting part is not the queue — it is designing Task so that a new kind of work can be added without touching the scheduler.

The design uses one abstract type, one concrete struct per kind of task, and multiple dispatch for the part that differs: how long a task takes and what it does. The scheduler itself never inspects a type — dispatch does the work.

# scheduler.jl — a priority scheduler built on abstract types
module Scheduler

export Task, Job, Review, enqueue!, run_next!, pending, by_priority

"""
    Task

Everything schedulable. Subtypes must implement `cost(task)` (a duration in
abstract units) and `execute(task)` (the work itself).
"""
abstract type Task end

"""A unit of work that takes `minutes` and reports `units` of output."""
struct Job <: Task
    name     :: String
    minutes  :: Float64
    units    :: Int
    priority :: Int          # lower number = more urgent
end

"""Time-boxed work: it takes at most `minutes`, whatever it produces."""
struct Review <: Task
    name     :: String
    minutes  :: Float64
    priority :: Int
end

# The two functions the abstract type requires. Adding a Task subtype means
# adding methods here — no change to the scheduler below.
cost(t::Job)    = t.minutes
cost(t::Review) = t.minutes

execute(t::Job)    = (t.units * 10, "produced $(t.units) units")
execute(t::Review) = (0, "reviewed $(t.name)")

"""
    Schedule

A priority queue kept sorted on insertion by `(priority, cost)`. A real
scheduler would use a heap; a sorted vector is easier to read and fast
enough for thousands of tasks.
"""
mutable struct Schedule
    tasks :: Vector{Task}
end
Schedule() = Schedule(Task[])

function enqueue!(s::Schedule, t::Task)
    push!(s.tasks, t)
    sort!(s.tasks, by = t -> (t.priority, cost(t)))
    return s
end

"Remove and run the most urgent task."
function run_next!(s::Schedule)
    isempty(s.tasks) && return nothing
    t = popfirst!(s.tasks)                 # the vector is already sorted
    value, description = execute(t)
    println(rpad(t.name, 18), rpad(string(cost(t)), 8), description)
    return t
end

pending(s::Schedule)     = length(s.tasks)
by_priority(s::Schedule) = sort(s.tasks, by = t -> t.priority)

end # module

# ---------------------------------------------------------------
if abspath(PROGRAM_FILE) == @__FILE__
    using .Scheduler

    s = Scheduler.Schedule()
    Scheduler.enqueue!(s, Scheduler.Job("build report", 30.0, 5, 2))
    Scheduler.enqueue!(s, Scheduler.Job("deploy hotfix", 10.0, 1, 1))
    Scheduler.enqueue!(s, Scheduler.Review("design doc", 45.0, 3))
    Scheduler.enqueue!(s, Scheduler.Job("refactor", 120.0, 20, 4))

    while (t = Scheduler.run_next!(s)) !== nothing; end
    @info "queue drained" remaining = Scheduler.pending(s)
end

The instructive line is sort!(s.tasks, by = t -> (t.priority, cost(t))): it calls cost on a Task whose concrete type it does not know, and dispatch does the right thing for each subtype. Adding a Meeting type later means writing two methods and changing nothing else — that is the payoff of the abstract-type-first design.

Challenge

Add three things. First, a Meeting subtype whose execute returns a named tuple rather than a string, which forces you to decide whether execute should have a common return type at all — write down what you decide and why. Second, a deadline field on every task and a scheduler that refuses to run anything past its deadline, returning the skipped tasks instead of dropping them silently. Third, replace the sorted vector with a binary heap and record how the running time changes as the queue grows from a hundred to a hundred thousand tasks. The third part is the reason the second exists: a correct answer that takes an hour is not a scheduler.

Project 3 — Epidemic Simulator

The last project is numerical. It solves the SIR model — the same equations the scientific computing lesson used — with an integrator written by hand, so that the connection between the mathematics and the code is visible on every line.

The design is a fourth-order Runge-Kutta step applied in a loop, with the results collected into vectors and written to a CSV file that any spreadsheet or plotting tool can read. No packages, no black boxes.

# sir_sim.jl — SIR model integrated with a hand-written RK4 stepper
module SIR

export Params, sir_deriv, rk4_step, simulate, write_csv

"""Epidemic parameters, with a derived reproduction number."""
struct Params
    β :: Float64        # transmission rate per contact per day
    γ :: Float64        # recovery rate per day
    population :: Float64
end

R0(p::Params) = p.β / p.γ            # basic reproduction number

"""
    sir_deriv(u, p) -> Vector{Float64}

The right-hand side: du/dt for (susceptible, infected, recovered).
N is conserved exactly by construction — the first sanity check.
"""
function sir_deriv(u::AbstractVector, p::Params)
    S, I, R = u
    N = S + I + R
    new_infections = p.β * S * I / N
    return [-new_infections,
             new_infections - p.γ * I,
             p.γ * I]
end

"""
    rk4_step(f, u, p, h) -> Vector{Float64}

One classical Runge-Kutta 4 step of size `h`. Written out explicitly rather
than in a loop over the coefficients, so the method is readable next to the
textbook it came from.
"""
function rk4_step(f, u::AbstractVector, p, h::Float64)
    k1 = f(u,              p)
    k2 = f(u .+ h/2 .* k1, p)
    k3 = f(u .+ h/2 .* k2, p)
    k4 = f(u .+ h   .* k3, p)
    return u .+ (h/6) .* (k1 .+ 2 .* k2 .+ 2 .* k3 .+ k4)
end

"""
    simulate(p; days, h) -> NamedTuple

Integrate from (N-1, 1, 0) for `days` days with step `h`, keeping one
sample per day so the output has a predictable size.
"""
function simulate(p::Params; days::Int = 160, h::Float64 = 0.1)
    steps_per_day = round(Int, 1 / h)
    u = [p.population - 1.0, 1.0, 0.0]        # one infected individual
    t = 0.0
    times  = Float64[0.0]
    states = Vector{Vector{Float64}}([copy(u)])

    for day in 1:days
        for _ in 1:steps_per_day
            u = rk4_step(sir_deriv, u, p, h)
            t += h
        end
        push!(times, t)
        push!(states, copy(u))                # copy: u is reused next step
    end

    peak_index = argmax(states[i][2] for i in eachindex(states))
    return (times = times, states = states, R0 = R0(p),
            peak_day = times[peak_index], peak_infected = states[peak_index][2])
end

"""Write day,S,I,R rows so any spreadsheet can read the result."""
function write_csv(path::AbstractString, result)
    open(path, "w") do io
        println(io, "day,susceptible,infected,recovered")
        for (t, u) in zip(result.times, result.states)
            println(io, round(t, digits = 2), ",", join(round.(u, digits = 4), ","))
        end
    end
    return path
end

end # module

# ---------------------------------------------------------------
if abspath(PROGRAM_FILE) == @__FILE__
    using .SIR
    p = SIR.Params(0.30, 0.10, 1000.0)
    result = SIR.simulate(p; days = 160, h = 0.1)
    SIR.write_csv("sir.csv", result)
    @info "simulation complete" R0 = result.R0 peak_day = result.peak_day peak_infected = round(result.peak_infected)
end

The invariants worth checking immediately are the ones that catch real numerical bugs: S + I + R must stay equal to population at every step, the infected curve must be smooth, and the infected count must never go negative. Run it once with h = 0.5 and once with h = 0.1 and compare the peak — the difference is the integration error, and seeing it is the point of writing your own stepper.

Challenge

Four extensions, in increasing difficulty. Add a vaccination term that moves people from susceptible to recovered at a constant daily rate, and find the rate at which the peak disappears. Add a step-size controller that halves h and retries when the difference between two half-steps exceeds a tolerance — the poor man's adaptive solver. Add a @testset that asserts conservation of N to within 1e-9 and the positivity of every compartment at every sample. Then replace rk4_step with a stiffly-stable implicit method and explain, in a comment, why the implicit form matters when γ is a thousand times larger than β.

What these three projects have in common. Each one separates a small, testable core (a module) from the script that drives it, uses the standard library instead of reaching for a package, and states its invariants in a comment or a test. That structure — core module, thin script, explicit invariants — is the shape of every well-behaved Julia program, from a ten-line utility to a simulation that runs for a week.

For the libraries that would replace these hand-written cores in production, continue to References & Playgrounds.