Data Definitions

Registers are scarce, so real programs keep most of their data in memory. This lesson is the assembler's vocabulary for describing that data: how many bytes, what initial values, and whether it may be written.

Defining Data

Data is described by a pseudo-instruction: it is not a CPU instruction at all, but a directive that tells the assembler to emit bytes. Every definition has a label (a name for the address) and a size-specific directive.

Bytes: db

db means "define byte". You can list as many values as you like, and a string in quotes is expanded into one byte per character.

section .data
    answer      db  42                  ; one byte: 0x2A
    newline     db  10                  ; '\n'
    zero        db  0                   ; a NUL terminator
    letters     db  'a', 'b', 'c'        ; three bytes: 61 62 63
    greeting    db  "Hello", 0          ; six bytes: the text plus a 0
    flags       db  0b10110001          ; binary and hex literals work too

Words, Dwords, and Qwords

The size prefix follows the same pattern: word = 2 bytes, dword = 4 bytes, qword = 8 bytes. Choose the smallest size that holds your values.

section .data
    count16  dw  65535                 ; 2 bytes, max unsigned 16-bit
    counter  dd  1000000               ; 4 bytes — the usual "int"
    address  dq  0x00007fff5fbff000    ; 8 bytes — a 64-bit pointer
    ratio    dq  3.14159               ; 8 bytes holding a float bit pattern

A dd holding arithmetic values is the safe default. Reserve dq for pointers and 64-bit counters, and remember that a dq costs eight bytes of cache for every value.

Repeating with times

times repeats a definition, which is the concise way to build tables, padding, and buffers.

section .data
    zeros    times 16 db 0             ; 16 zero bytes
    squares  times 5  dd 0, 1, 4, 9, 16 ; WRONG: times repeats the whole list
                                        ; -> use separate definitions instead

section .bss
    buffer   resb 4096                 ; reserve 4096 bytes, no initial value
Careful: times n db v1, v2 emits the entire list n times. To build a lookup table, write the values out explicitly or generate them with an assembler macro, as shown in Macros & Directives.

Strings and Characters

Assembly has no string type. A string is a sequence of bytes with a label pointing at the first one, and it is your job to remember how long it is. Two conventions dominate:

  • Length-prefixed / length-counted — you keep the byte count in a register or constant. This is what Linux system calls expect.
  • NUL-terminated — a zero byte marks the end. This is what C library functions expect.
section .rodata
    ; Linux syscalls want: pointer + explicit length
    msg      db  "Hello", 10
    MSG_LEN  equ $ - msg                ; 6

    ; C functions want: pointer + terminating zero
    cstr     db  "Hello", 0

Mixing the two conventions is a classic bug: passing a length-counted string to printf prints whatever random bytes follow it until a zero appears. Always check which convention the function you are calling expects.

Uninitialized Data: the .bss Section

Large buffers do not belong in .data. Writing times 65536 db 0 would put 64 KB of zeros inside your executable file. The .bss section solves this: it reserves space that the operating system fills with zeros at load time, costing nothing on disk.

section .bss
    read_buf  resb 4096       ; 4 KB of zeroed memory, free on disk
    total     resq 1          ; one 8-byte slot
    tmp       resd 16         ; an array of sixteen 32-bit slots

The trade-off is that .bss cannot hold initial values other than zero — if you need 1 or "hello", it must go in .data or .rodata.

Reading and Writing Memory

Sizes Must Match — or Be Stated

When one operand is a register, the assembler knows the size and everything is unambiguous. When both operands are memory or a constant, the assembler cannot guess, and you must say how many bytes to move with a size qualifier.

section .data
    byte_val  db 0
    word_val  dw 0
    dword_val dd 0

section .text
    mov  al, [byte_val]           ; size comes from al (8 bits) — fine
    mov  [byte_val], al           ; still fine: al supplies the size

    mov  byte [byte_val], 7       ; constant -> memory: MUST give the size
    mov  word [word_val], 7
    mov  dword [dword_val], 7

    ; mov  [byte_val], 7          ; ERROR: operation size not specified
    ; mov  eax, [word_val]        ; ERROR: operand size mismatch

Size mismatches are the single most common assembly error message, and the fix is always the same: name the size, or route the value through a register of the right width.

mov vs lea

These two look similar and mean completely different things. mov copies the contents at an address; lea computes the address itself and never touches memory.

section .data
    msg  db "hello", 0

section .text
    mov  rax, [msg]        ; rax = 'h' 'e' 'l' 'l' + 4 zero bytes  (the DATA)
    lea  rbx, [msg]        ; rbx = the ADDRESS of msg             (the POINTER)

    ; lea can also do pure arithmetic with no memory access at all:
    lea  rcx, [rdi + rsi*4 + 8]   ; rcx = rdi + rsi*4 + 8

Because lea can compute base-plus-index-plus-displacement in one instruction, compilers use it for ordinary arithmetic as much as for addresses. The next lesson, Addressing Modes, is devoted to that expression syntax.

Summary

  • db, dw, dd, dq define 1-, 2-, 4-, and 8-byte values; times n repeats a definition.
  • Initialized data lives in .data (or .rodata if read-only); uninitialized space is reserved in .bss with resb/resq.
  • A string is just bytes with a label; you must keep its length yourself, either as an equ or with a terminating zero.
  • Give a size qualifier (byte, word, dword, qword) whenever both operands are memory or a constant.
  • mov reads memory; lea computes an address without reading anything.

Next: Addressing Modes — the full expression syntax that turns registers into addresses.