Data Definitions
Defining Data
Data is described by a pseudo-instruction: it is not a CPU instruction at all, but a directive that tells the assembler to emit bytes. Every definition has a label (a name for the address) and a size-specific directive.
Bytes: db
db means "define byte". You can list as many values as you like, and a string in quotes is expanded into one byte per character.
section .data
answer db 42 ; one byte: 0x2A
newline db 10 ; '\n'
zero db 0 ; a NUL terminator
letters db 'a', 'b', 'c' ; three bytes: 61 62 63
greeting db "Hello", 0 ; six bytes: the text plus a 0
flags db 0b10110001 ; binary and hex literals work too
Words, Dwords, and Qwords
The size prefix follows the same pattern: word = 2 bytes, dword = 4 bytes, qword = 8 bytes. Choose the smallest size that holds your values.
section .data
count16 dw 65535 ; 2 bytes, max unsigned 16-bit
counter dd 1000000 ; 4 bytes — the usual "int"
address dq 0x00007fff5fbff000 ; 8 bytes — a 64-bit pointer
ratio dq 3.14159 ; 8 bytes holding a float bit pattern
A dd holding arithmetic values is the safe default. Reserve dq for pointers and 64-bit counters, and remember that a dq costs eight bytes of cache for every value.
Repeating with times
times repeats a definition, which is the concise way to build tables, padding, and buffers.
section .data
zeros times 16 db 0 ; 16 zero bytes
squares times 5 dd 0, 1, 4, 9, 16 ; WRONG: times repeats the whole list
; -> use separate definitions instead
section .bss
buffer resb 4096 ; reserve 4096 bytes, no initial value
times n db v1, v2 emits the entire list n times. To build a lookup table, write the values out explicitly or generate them with an assembler macro, as shown in Macros & Directives.
Strings and Characters
Assembly has no string type. A string is a sequence of bytes with a label pointing at the first one, and it is your job to remember how long it is. Two conventions dominate:
- Length-prefixed / length-counted — you keep the byte count in a register or constant. This is what Linux system calls expect.
- NUL-terminated — a zero byte marks the end. This is what C library functions expect.
section .rodata
; Linux syscalls want: pointer + explicit length
msg db "Hello", 10
MSG_LEN equ $ - msg ; 6
; C functions want: pointer + terminating zero
cstr db "Hello", 0
Mixing the two conventions is a classic bug: passing a length-counted string to printf prints whatever random bytes follow it until a zero appears. Always check which convention the function you are calling expects.
Uninitialized Data: the .bss Section
Large buffers do not belong in .data. Writing times 65536 db 0 would put 64 KB of zeros inside your executable file. The .bss section solves this: it reserves space that the operating system fills with zeros at load time, costing nothing on disk.
section .bss
read_buf resb 4096 ; 4 KB of zeroed memory, free on disk
total resq 1 ; one 8-byte slot
tmp resd 16 ; an array of sixteen 32-bit slots
The trade-off is that .bss cannot hold initial values other than zero — if you need 1 or "hello", it must go in .data or .rodata.
Reading and Writing Memory
Sizes Must Match — or Be Stated
When one operand is a register, the assembler knows the size and everything is unambiguous. When both operands are memory or a constant, the assembler cannot guess, and you must say how many bytes to move with a size qualifier.
section .data
byte_val db 0
word_val dw 0
dword_val dd 0
section .text
mov al, [byte_val] ; size comes from al (8 bits) — fine
mov [byte_val], al ; still fine: al supplies the size
mov byte [byte_val], 7 ; constant -> memory: MUST give the size
mov word [word_val], 7
mov dword [dword_val], 7
; mov [byte_val], 7 ; ERROR: operation size not specified
; mov eax, [word_val] ; ERROR: operand size mismatch
Size mismatches are the single most common assembly error message, and the fix is always the same: name the size, or route the value through a register of the right width.
mov vs lea
These two look similar and mean completely different things. mov copies the contents at an address; lea computes the address itself and never touches memory.
section .data
msg db "hello", 0
section .text
mov rax, [msg] ; rax = 'h' 'e' 'l' 'l' + 4 zero bytes (the DATA)
lea rbx, [msg] ; rbx = the ADDRESS of msg (the POINTER)
; lea can also do pure arithmetic with no memory access at all:
lea rcx, [rdi + rsi*4 + 8] ; rcx = rdi + rsi*4 + 8
Because lea can compute base-plus-index-plus-displacement in one instruction, compilers use it for ordinary arithmetic as much as for addresses. The next lesson, Addressing Modes, is devoted to that expression syntax.
Summary
db,dw,dd,dqdefine 1-, 2-, 4-, and 8-byte values;times nrepeats a definition.- Initialized data lives in
.data(or.rodataif read-only); uninitialized space is reserved in.bsswithresb/resq. - A string is just bytes with a label; you must keep its length yourself, either as an
equor with a terminating zero. - Give a size qualifier (
byte,word,dword,qword) whenever both operands are memory or a constant. movreads memory;leacomputes an address without reading anything.
Next: Addressing Modes — the full expression syntax that turns registers into addresses.