Study Projects
Each project is complete as shown and takes about an hour to write and understand. The challenge at the end of each is deliberately open: it is the version of the project that you would meet in real work, where the requirements arrive as a paragraph rather than as a specification.
Project 1 — Word Frequency Counter
Every programmer writes this program once. It is the smallest project that exercises files, strings, dictionaries, sorting, and formatting together — and the smallest one where the obvious implementation is also the slow one if you let it re-read the file.
The design: read the file once, normalise each word, count with a Dict, then sort the entries. Nothing here needs a package, and every step maps onto a lesson from earlier in the track.
# word_frequency.jl — count words in a text file
module WordFrequency
export count_words, top_words, report
"""
count_words(path) -> Dict{String,Int}
Read the file at `path` once and return word => count.
A word is a run of letters, digits, apostrophes or hyphens; everything
else is a separator. Matching is case-insensitive.
"""
function count_words(path::AbstractString)
isfile(path) || throw(ArgumentError("no such file: $path"))
counts = Dict{String,Int}()
for line in eachline(path) # streams: no whole-file copy
for raw in split(line)
word = normalise(raw)
isempty(word) && continue # punctuation-only token
counts[word] = get(counts, word, 0) + 1
end
end
return counts
end
"""
Lowercase a token and strip leading/trailing punctuation.
Everything inside the token is kept, so "state-of-the-art" stays one word.
"""
function normalise(token::AbstractString)
# keep letters, digits, apostrophes and hyphens; drop the rest
cleaned = filter(c -> isletter(c) || isdigit(c) || c in ('\'', '-'), lowercase(token))
return strip(cleaned, ['\'', '-']) # "word." -> "word", "'word'" -> "word"
end
"""
top_words(counts, n) -> Vector{Pair{String,Int}}
The `n` most frequent words, most frequent first. Ties break alphabetically
so the output is stable between runs — important for tests and for diffs.
"""
function top_words(counts::AbstractDict, n::Integer = 10)
entries = collect(counts)
sort!(entries, by = e -> (-e.second, e.first))
return entries[1:min(n, length(entries))]
end
"""Render a report with counts and a simple bar chart."""
function report(path::AbstractString; top::Integer = 10)
counts = count_words(path)
words = top_words(counts, top)
total = sum(values(counts))
unique_words = length(counts)
widest = isempty(words) ? 0 : maximum(length(first(w)) for w in words)
println("file : ", path)
println("tokens : ", total)
println("distinct words : ", unique_words)
println("type/token ratio: ", round(unique_words / max(total, 1), digits = 3))
println()
for (word, count) in words
bar = repeat("#", min(count, 40)) # cap the bar at terminal width
println(rpad(word, widest), " ", lpad(count, 6), " ", bar)
end
return counts
end
end # module
# ---------------------------------------------------------------
if abspath(PROGRAM_FILE) == @__FILE__ # only when run as a script
using .WordFrequency
isempty(ARGS) && (println(stderr, "usage: julia word_frequency.jl FILE [N]"); exit(2))
n = length(ARGS) >= 2 ? parse(Int, ARGS[2]) : 10
WordFrequency.report(ARGS[1]; top = n)
end
Three details carry the lesson. eachline(path) streams the file instead of loading it, so the program survives a 500 MB log. get(counts, word, 0) + 1 is the idiomatic counter — a single hash lookup, no branch. And sorting by a tuple (-count, word) gives the stable ordering that makes the output testable.
Challenge
Extend the program in three steps, in this order. First, add a --min-length option that ignores words shorter than a given length — the option parser belongs in the script block, not in the module. Second, make it read standard input when the path is -, so it can be used in a pipeline: cat book.txt | julia word_frequency.jl -. Third, add a @testset that checks normalise against a table of awkward cases ("End.", "'quoted'", "co-op", "123") and checks that a tie in counts produces alphabetical order. When the third step is easy, the design is right.
Project 2 — Priority Task Scheduler
The second project is about types. A scheduler holds heterogeneous work items, orders them by priority and deadline, and runs them. The interesting part is not the queue — it is designing Task so that a new kind of work can be added without touching the scheduler.
The design uses one abstract type, one concrete struct per kind of task, and multiple dispatch for the part that differs: how long a task takes and what it does. The scheduler itself never inspects a type — dispatch does the work.
# scheduler.jl — a priority scheduler built on abstract types
module Scheduler
export Task, Job, Review, enqueue!, run_next!, pending, by_priority
"""
Task
Everything schedulable. Subtypes must implement `cost(task)` (a duration in
abstract units) and `execute(task)` (the work itself).
"""
abstract type Task end
"""A unit of work that takes `minutes` and reports `units` of output."""
struct Job <: Task
name :: String
minutes :: Float64
units :: Int
priority :: Int # lower number = more urgent
end
"""Time-boxed work: it takes at most `minutes`, whatever it produces."""
struct Review <: Task
name :: String
minutes :: Float64
priority :: Int
end
# The two functions the abstract type requires. Adding a Task subtype means
# adding methods here — no change to the scheduler below.
cost(t::Job) = t.minutes
cost(t::Review) = t.minutes
execute(t::Job) = (t.units * 10, "produced $(t.units) units")
execute(t::Review) = (0, "reviewed $(t.name)")
"""
Schedule
A priority queue kept sorted on insertion by `(priority, cost)`. A real
scheduler would use a heap; a sorted vector is easier to read and fast
enough for thousands of tasks.
"""
mutable struct Schedule
tasks :: Vector{Task}
end
Schedule() = Schedule(Task[])
function enqueue!(s::Schedule, t::Task)
push!(s.tasks, t)
sort!(s.tasks, by = t -> (t.priority, cost(t)))
return s
end
"Remove and run the most urgent task."
function run_next!(s::Schedule)
isempty(s.tasks) && return nothing
t = popfirst!(s.tasks) # the vector is already sorted
value, description = execute(t)
println(rpad(t.name, 18), rpad(string(cost(t)), 8), description)
return t
end
pending(s::Schedule) = length(s.tasks)
by_priority(s::Schedule) = sort(s.tasks, by = t -> t.priority)
end # module
# ---------------------------------------------------------------
if abspath(PROGRAM_FILE) == @__FILE__
using .Scheduler
s = Scheduler.Schedule()
Scheduler.enqueue!(s, Scheduler.Job("build report", 30.0, 5, 2))
Scheduler.enqueue!(s, Scheduler.Job("deploy hotfix", 10.0, 1, 1))
Scheduler.enqueue!(s, Scheduler.Review("design doc", 45.0, 3))
Scheduler.enqueue!(s, Scheduler.Job("refactor", 120.0, 20, 4))
while (t = Scheduler.run_next!(s)) !== nothing; end
@info "queue drained" remaining = Scheduler.pending(s)
end
The instructive line is sort!(s.tasks, by = t -> (t.priority, cost(t))): it calls cost on a Task whose concrete type it does not know, and dispatch does the right thing for each subtype. Adding a Meeting type later means writing two methods and changing nothing else — that is the payoff of the abstract-type-first design.
Challenge
Add three things. First, a Meeting subtype whose execute returns a named tuple rather than a string, which forces you to decide whether execute should have a common return type at all — write down what you decide and why. Second, a deadline field on every task and a scheduler that refuses to run anything past its deadline, returning the skipped tasks instead of dropping them silently. Third, replace the sorted vector with a binary heap and record how the running time changes as the queue grows from a hundred to a hundred thousand tasks. The third part is the reason the second exists: a correct answer that takes an hour is not a scheduler.
Project 3 — Epidemic Simulator
The last project is numerical. It solves the SIR model — the same equations the scientific computing lesson used — with an integrator written by hand, so that the connection between the mathematics and the code is visible on every line.
The design is a fourth-order Runge-Kutta step applied in a loop, with the results collected into vectors and written to a CSV file that any spreadsheet or plotting tool can read. No packages, no black boxes.
# sir_sim.jl — SIR model integrated with a hand-written RK4 stepper
module SIR
export Params, sir_deriv, rk4_step, simulate, write_csv
"""Epidemic parameters, with a derived reproduction number."""
struct Params
β :: Float64 # transmission rate per contact per day
γ :: Float64 # recovery rate per day
population :: Float64
end
R0(p::Params) = p.β / p.γ # basic reproduction number
"""
sir_deriv(u, p) -> Vector{Float64}
The right-hand side: du/dt for (susceptible, infected, recovered).
N is conserved exactly by construction — the first sanity check.
"""
function sir_deriv(u::AbstractVector, p::Params)
S, I, R = u
N = S + I + R
new_infections = p.β * S * I / N
return [-new_infections,
new_infections - p.γ * I,
p.γ * I]
end
"""
rk4_step(f, u, p, h) -> Vector{Float64}
One classical Runge-Kutta 4 step of size `h`. Written out explicitly rather
than in a loop over the coefficients, so the method is readable next to the
textbook it came from.
"""
function rk4_step(f, u::AbstractVector, p, h::Float64)
k1 = f(u, p)
k2 = f(u .+ h/2 .* k1, p)
k3 = f(u .+ h/2 .* k2, p)
k4 = f(u .+ h .* k3, p)
return u .+ (h/6) .* (k1 .+ 2 .* k2 .+ 2 .* k3 .+ k4)
end
"""
simulate(p; days, h) -> NamedTuple
Integrate from (N-1, 1, 0) for `days` days with step `h`, keeping one
sample per day so the output has a predictable size.
"""
function simulate(p::Params; days::Int = 160, h::Float64 = 0.1)
steps_per_day = round(Int, 1 / h)
u = [p.population - 1.0, 1.0, 0.0] # one infected individual
t = 0.0
times = Float64[0.0]
states = Vector{Vector{Float64}}([copy(u)])
for day in 1:days
for _ in 1:steps_per_day
u = rk4_step(sir_deriv, u, p, h)
t += h
end
push!(times, t)
push!(states, copy(u)) # copy: u is reused next step
end
peak_index = argmax(states[i][2] for i in eachindex(states))
return (times = times, states = states, R0 = R0(p),
peak_day = times[peak_index], peak_infected = states[peak_index][2])
end
"""Write day,S,I,R rows so any spreadsheet can read the result."""
function write_csv(path::AbstractString, result)
open(path, "w") do io
println(io, "day,susceptible,infected,recovered")
for (t, u) in zip(result.times, result.states)
println(io, round(t, digits = 2), ",", join(round.(u, digits = 4), ","))
end
end
return path
end
end # module
# ---------------------------------------------------------------
if abspath(PROGRAM_FILE) == @__FILE__
using .SIR
p = SIR.Params(0.30, 0.10, 1000.0)
result = SIR.simulate(p; days = 160, h = 0.1)
SIR.write_csv("sir.csv", result)
@info "simulation complete" R0 = result.R0 peak_day = result.peak_day peak_infected = round(result.peak_infected)
end
The invariants worth checking immediately are the ones that catch real numerical bugs: S + I + R must stay equal to population at every step, the infected curve must be smooth, and the infected count must never go negative. Run it once with h = 0.5 and once with h = 0.1 and compare the peak — the difference is the integration error, and seeing it is the point of writing your own stepper.
Challenge
Four extensions, in increasing difficulty. Add a vaccination term that moves people from susceptible to recovered at a constant daily rate, and find the rate at which the peak disappears. Add a step-size controller that halves h and retries when the difference between two half-steps exceeds a tolerance — the poor man's adaptive solver. Add a @testset that asserts conservation of N to within 1e-9 and the positivity of every compartment at every sample. Then replace rk4_step with a stiffly-stable implicit method and explain, in a comment, why the implicit form matters when γ is a thousand times larger than β.
For the libraries that would replace these hand-written cores in production, continue to References & Playgrounds.