Arithmetic
Add, Subtract, and the Flags
Arithmetic instructions do two things at once: they produce a result and they record facts about that result in the flags register. Ignore the flags and you are programming with one hand tied behind your back.
add and sub
mov rax, 50
add rax, 8 ; rax = 58
sub rax, 20 ; rax = 38 (sub is exactly add with the sign flipped)
add [total], rax ; memory can be a destination: total += rax
add rax, [total] ; ...or a source: rax += total
Both set the full flag set. After any add or sub you may branch on zero (jz), on sign (js), on unsigned carry (jc), or on signed overflow (jo) — at no extra cost, because the flags are already there.
inc and dec
inc rcx ; rcx += 1 — one byte shorter than `add rcx, 1`
dec rbx ; rbx -= 1
inc and dec set every flag except the carry flag. That one omission matters: if you are using the carry flag to hold state across a loop, an inc inside the loop will not disturb it, whereas add would.
neg
mov rax, 42
neg rax ; rax = -42 (invert all bits, then add 1)
Negation works by two's complement, which is why subtraction needs no separate hardware: a - b is computed as a + neg(b). The negation of the most negative value overflows and sets OF; it is the one case that stays negative.
Carrying Across 64 Bits: adc and sbb
To add numbers wider than 64 bits, process them one word at a time from the least significant end, and let the carry propagate with adc ("add with carry") and sbb ("subtract with borrow").
; Add two 128-bit numbers stored as [low, high] pairs
mov rax, [a_low]
add rax, [b_low] ; add the low words; sets CF if it carried out
mov [r_low], rax
mov rax, [a_high]
adc rax, [b_high] ; add the high words PLUS the carry from the low add
mov [r_high], rax
; sbb does the same for subtraction, subtracting the borrow.
This pattern is how every big-integer library, checksum, and cryptographic routine handles numbers larger than the register width.
Multiplication and Division
These instructions break the tidy two-operand pattern, because the product of two n-bit numbers needs 2n bits, and the dividend of an n-bit result needs an input twice that wide.
Multiplication with imul
The modern two-operand imul is what you want almost every time: it multiplies two registers and stores the low half of the product in the destination.
mov rax, 12
mov rbx, 8
imul rax, rbx ; rax = rax * rbx = 96 (low 64 bits of the product)
imul rax, rbx, 4 ; three-operand form: rax = rbx * 4 (imm must be constant)
; For a full 128-bit product you need the one-operand form, which uses rdx:rax
mov rax, 0x7FFFFFFFFFFFFFFF
mov rbx, 2
imul rbx ; rdx:rax = rax * rbx, 128 bits total
; OF=1 if the high half is not just the sign extension
The one-operand form is only needed for arbitrary-precision work. For ordinary 64-bit values, two-operand imul is shorter and faster.
Division with idiv
Division is the most error-prone instruction in the whole set: the dividend is always held in rdx:rax, and the quotient and remainder come back in two fixed places.
; 100 / 7 (signed)
mov rax, 100 ; rax = dividend (low half)
cqo ; sign-extend rax into rdx:rax <-- REQUIRED before idiv
mov rcx, 7 ; rcx = divisor
idiv rcx ; rax = quotient = 14
; rdx = remainder = 2
; Unsigned division behaves the same but uses the other instructions:
mov rax, 100
xor rdx, rdx ; zero-extend rax into rdx:rax <-- use xor, not cqo
mov rcx, 7
div rcx ; rax = quotient, rdx = remainder
SIGFPE and kills the process; and forgetting cqo (signed) or xor rdx, rdx (unsigned) leaves garbage in rdx, which almost always produces a divide-overflow fault. Both are in the top five bugs of every assembly beginner, so check them first when your program dies with Floating point exception while doing integer maths.
Modulo Without Division
When the divisor is a power of two, the modulo is just a mask — no division at all:
; rax % 8, for a NON-NEGATIVE rax
and rax, 7 ; keep the low three bits — identical to rax % 8
; (for signed values, negative inputs need the correction term first)
A Note on Floating Point
Everything above is integer arithmetic. Floating-point numbers are handled by a completely separate register file and a separate instruction set: x87 (the historical stack-based unit) and SSE/AVX (the modern one).
; Single-precision (float) addition with SSE
movss xmm0, [a] ; load 4 bytes into a vector register
movss xmm1, [b]
addss xmm0, xmm1 ; xmm0 = a + b
movss [result], xmm0 ; store it back
; Double precision uses the sd (scalar double) variants:
; movsd / addsd / subsd / mulsd / divsd
The XMM registers are 128 bits wide and hold several numbers at once, which is the basis of SIMD. Floating-point programming deserves its own treatment, but for now know that it is a different world: different registers, different instructions, and different rounding rules.
Handling Overflow Correctly
There is no automatic error, no exception, and no warning: an arithmetic instruction overflow sets a flag and keeps going. If you do not check the flag, you get a wrong answer with no indication at all.
; Safe addition with an explicit overflow check
mov rax, a
add rax, b
jo overflow_error ; jo = "jump if overflow" (signed)
; ... continue with a valid result ...
overflow_error:
; handle it: return an error code, clamp, or abort
mov rax, 60 ; exit
mov rdi, 1 ; status 1 = failure
syscall
Two jump mnemonics do the checks:
jo/jno— jump if the signed overflow flag is set / clear.jc/jnc— jump if the unsigned carry flag is set / clear.jcandjbare the same instruction.
The Wider-Type Remedy
In practice, most code avoids the flag entirely by computing in a wider type and range-checking the result. The same trick works in assembly, and it is what compilers do for arithmetic they cannot prove safe.
; Add two 32-bit values without risking a 32-bit overflow:
mov eax, a ; zero-extends into rax
add rax, b ; 64-bit addition cannot overflow here
cmp rax, 0x7FFFFFFF ; compare against the largest signed 32-bit value
ja overflow_error ; result does not fit -> handle it
; rax now holds a correct 32-bit result
This is why C's int arithmetic and assembly arithmetic behave differently out of the box: the compiler inserts the range check for you, and hand-written assembly does not.
Summary
add,sub,inc,dec, andnegare the basic arithmetic instructions; all set the flags.- Use
adc/sbbto chain additions across 64-bit words using the carry flag. - Multiplication and division use
rax(andrdx), because their results are twice as wide as their inputs. idivrequirescqofirst and traps (SIGFPE) if the divisor is zero.- Never assume arithmetic is safe — check
jo/jcor compute in a wider type.
Next: Control Flow — turning those flags into decisions and loops.