Elixir: Strings, Binaries & Sigils

In Elixir, a string is a UTF-8 binary, and a binary is a bitstring whose bit count is divisible by eight. That single sentence unlocks binary pattern matching — one of the language's superpowers for protocol and file parsing.

Strings Are UTF-8 Binaries

s = "héllo"            # a binary: UTF-8 encoded bytes
byte_size(s)           #=> 6   — "é" is 2 bytes in UTF-8
String.length(s)       #=> 5   — counts CHARACTERS (code points)
String.upcase(s)       #=> "HÉLLO" — Unicode-aware

# Graphemes: user-perceived characters can span several code points.
String.length("a\u0301")        #=> 2 code points...
String.length(String.graphemes("a\u0301") |> hd())  # ...one grapheme

Rule of thumb: String functions are Unicode-aware and binary-based; the :binary module (from Erlang) works on raw bytes; Bitwise handles bit operations. Mixing them up is a classic source of off-by-one bugs on non-ASCII input.

Bitstring Syntax

<<>> builds and matches raw bits. Each segment declares its size and type — size: bits, unit: steps, type integer/binary/float, and utf8 for direct string embedding.

<<1, 2, 3>>                       # three bytes
<<1024::16>>                      # 16 bits holding 1024  => <<4, 0>>
<<3.14::float>>                   # a float packed into bytes
<<"é"::utf8>>                     # encode é explicitly
<<flag::1, payload::7>> = <<0b1000_0001>>   # flag=1, payload=1

Binary Pattern Matching in Practice

Parsing a wire protocol shows why binaries matter: the pattern is the grammar. This function decodes a length-prefixed frame and recurses — no manual byte counters, no off-by-one bugs.

# Frame format: [2-byte big-endian length][payload bytes]
defmodule Frame do
  def decode(<<length::16, payload::bytes-size(length), rest::binary>>) do
    # One clause matches a complete frame AND binds `rest` for the next one.
    [{:ok, payload} | decode(rest)]
  end

  def decode(<<>>), do: []                       # clean end of stream
  def decode(truncated), do: [{:error, {:truncated, byte_size(truncated)}}]
end

Frame.decode(<<5, "hello", 3, "abc", 2, "hi">>)
#=> [ok: "hello", ok: "abc", ok: "hi"]

The same technique parses timestamps, IP headers, PNG chunks — anything with a byte layout. When a format is bit-oriented rather than byte-oriented, only the sizes change.

Charlists and IO Data

'hello'          # charlist: a LIST of integer code points (Erlang heritage)
:io.format("~s~n", [~c"from charlist"])   # ~c sigil builds charlists

# IO data defers concatenation: build huge outputs without copies.
iodata = ["status=", "ok", "\n", ["rows=", "42"]]
IO.iodata_to_binary(iodata)   #=> "status=ok\nrows=42"

IO data matters at scale: logging frameworks and Phoenix responses accept it so that thousand-part responses are assembled with zero intermediate strings.

Sigils in Depth

Sigils from the syntax lesson generalize: any lowercase letter can become a custom sigil by defining sigil_x/2 in a module. The built-ins worth memorizing: ~s strings, ~w word lists, ~r regexes, ~D/~T/~N/~U date-time structs.

Practice

  1. Write a decoder for a header of [1-byte version][4-byte timestamp][payload] and return {:error, :bad_version} for unsupported versions via guards.
  2. Measure Enum.join versus IO data for 10,000 parts (use :timer.tc/1); note the difference.
  3. Explain the difference between byte_size/1 and String.length/1 using the string "café".

Next: Macros & Metaprogramming