Numbers & Memory

A CPU has no idea what a number means. It only has bits. Signedness, decimal points, and characters are conventions that you — the programmer — must apply consistently, because nothing below will check for you.

Number Bases

Three bases are used every day in assembly, and the reason is practical rather than mathematical: each one makes a different thing easy to see.

Binary — What the Hardware Sees

A wire is either at 0 V or 5 V. That is the entire vocabulary of the machine, so every number is a sequence of ones and zeros. An n-bit value holds 2n distinct patterns:

WidthBitsDistinct valuesx86-64 register
byte8256al
word1665 536ax
double word324 294 967 296eax
quad word6418 446 744 073 709 551 616rax

Hexadecimal — Shorthand for Bits

One hexadecimal digit is exactly four bits, so two hex digits are exactly one byte. That one-to-one correspondence is why hex is everywhere in debugging: 0x9F is 1001 1111, instantly, with no arithmetic.

DecHexBinDecHexBin
000000881000
110001991001
22001010A1010
33001111B1011
44010012C1100
55010113D1101
66011014E1110
77011115F1111

Converting to hex is mechanical: group the bits in fours from the right, then look each group up. To compute a value by hand: each hex digit is worth 16position, so 0x2A is 2×16 + 10 = 42.

Signed Numbers

Two's Complement and the Range Problem

Nothing in the hardware says a bit pattern is negative. The convention that makes negative numbers cheap is two's complement: the most significant bit carries a negative weight. For an 8-bit value, bit 7 is worth −128 instead of +128.

    ; 8-bit patterns interpreted as UNSIGNED and as SIGNED (two's complement):
    ;
    ;   bits      unsigned   signed
    ;   --------  --------   ------
    ;   0000 0000      0          0
    ;   0111 1111    127        127     <- largest signed positive
    ;   1000 0000    128       -128     <- most negative
    ;   1111 1111    255         -1

The range is therefore asymmetric: −2n−1 up to +2n−1−1. For 32-bit integers that is −2 147 483 648 to 2 147 483 647 — the familiar int limits of most languages.

Negation and Subtraction

To negate a number in two's complement, invert every bit and add one. That trick is why the CPU needs only an adder: a − b is computed as a + (−b).

    mov  al, 5          ; 0000 0101
    neg  al             ; invert and add 1 -> 1111 1011 = -5

    ;   ~5      = 1111 1010   (not al)
    ;   ~5 + 1  = 1111 1011   (neg al)  = -5

Overflow

Signed and unsigned arithmetic can both fail, but they fail in different ways, and the CPU reports each with its own flag:

  • Carry flag (CF) — set when an unsigned result does not fit (e.g. 255 + 1 in a byte).
  • Overflow flag (OF) — set when a signed result has the wrong sign (e.g. 127 + 1 in a byte = −128).
    mov  al, 127
    add  al, 1          ; result 1000 0000
                        ; CF = 0 (128 fits in an unsigned byte)
                        ; OF = 1 (127 + 1 overflows signed range: now -128)

This is exactly why you must know whether your data is signed before you choose a jump instruction. The next lesson on Control Flow shows that jg (signed) and ja (unsigned) read the same flags but reach opposite conclusions.

Bytes, Endianness, and Alignment

A byte is eight bits. Larger numbers occupy several consecutive bytes, and the CPU must agree on which byte holds the most significant part:

  • Little-endian — the least significant byte comes first in memory. This is what x86-64 and default-mode ARM use.
  • Big-endian — the most significant byte comes first. This is what network protocols and some RISC architectures use.
    ; The 32-bit value 0x12345678 stored at address `val`:
    ;
    ;   little-endian (x86-64):   78 56 34 12
    ;   big-endian:               12 34 56 78
    ;
    ; The number is the same; only the byte order in memory differs.

section .data
    val  dd  0x12345678        ; a 32-bit little-endian value

Endianness only becomes visible when you inspect memory byte by byte — in a hex dump, over a network, or inside a file format. It is the classic source of "the file is written backwards" bugs.

Alignment is the other memory convention: an 8-byte value is fastest when its address is a multiple of 8, a 4-byte value when its address is a multiple of 4, and so on. x86-64 tolerates misalignment with a small penalty; ARM and RISC-V may fault outright. NASM's align directive pads to a boundary so you never have to think about it:

section .data
    flag    db  1              ; 1 byte
    align   8                  ; pad with zeros up to the next multiple of 8
    total   dq  0              ; now guaranteed to be 8-byte aligned

Number Literals in NASM

Writing Numbers in Source

NASM accepts several spellings for the same value. Choose the one that matches the meaning: hexadecimal for bit patterns and addresses, binary when individual bits matter, decimal for everyday quantities.

    mov  eax, 255          ; decimal
    mov  eax, 0xFF         ; hexadecimal, 0x prefix      (= 255)
    mov  eax, 0ffh         ; hexadecimal, h suffix (leading 0 required)
    mov  eax, $FF          ; hexadecimal, $ prefix
    mov  eax, 0b11111111   ; binary, 0b prefix           (= 255)
    mov  eax, 0o377        ; octal, 0o prefix            (= 255)
    mov  eax, 377q         ; octal, q suffix
    mov  eax, 1_000_000    ; underscores group digits    (= 1000000)

    ; Character and string constants: the assembler substitutes the code point.
    mov  al,  'A'          ; al = 65  (ASCII code of 'A')
    mov  al,  'A' + 1      ; al = 66  — expressions are allowed

Expressions Are Evaluated at Assembly Time

NASM computes constant expressions while assembling, so arithmetic on literals costs nothing at run time. This is how string lengths and masks are written without magic numbers.

    BUFFER_SIZE equ 4096
    HALF        equ BUFFER_SIZE / 2           ; 2048
    MASK        equ (1 << 12) - 1            ; 4095 — low 12 bits set
    msg         db "hi", 10
    MSG_LEN     equ $ - msg                   ; 3

    mov  rdx, MSG_LEN                         ; no run-time work at all

Bitwise Operations

Because everything is bits, the bitwise instructions are not a curiosity — they replace multiplication, division, and modulo whenever the constant happens to be a power of two.

OperationInstructionMeaningTypical use
ANDand dst, src1 where both bits are 1Mask off (keep) selected bits
ORor dst, src1 where either bit is 1Set selected bits
XORxor dst, src1 where bits differToggle bits; clear a register
NOTnot dstInvert every bitOne's complement
Shift leftshl dst, nMove bits left, fill with 0Multiply by 2n
Shift rightshr dst, nMove bits right, fill with 0Unsigned divide by 2n
Arithmetic shift rightsar dst, nShift right, copy the sign bitSigned divide by 2n
    mov  eax, 0b00001111
    and  eax, 0b00000011   ; keep only the low two bits   -> 0b00000011 (3)
    or   eax, 0b00010000   ; set bit 4                    -> 0b00010011 (19)
    xor  eax, 0b00000011   ; toggle the low two bits      -> 0b00010000 (16)
    shl  eax, 3            ; multiply by 8                -> 128
    shr  eax, 1            ; unsigned divide by 2         -> 64

Summary

  • The CPU stores only bits; signedness and meaning are conventions you apply.
  • Two's complement makes subtraction the same hardware as addition, at the cost of an asymmetric range (−2n−1 … 2n−1−1).
  • Signed and unsigned comparisons use different jump instructions (jg vs ja) on the same flag bits.
  • Little-endian is the byte order on x86-64; alignment matters far more on ARM and RISC-V than on x86.
  • NASM evaluates constant expressions at assembly time — use equ instead of magic numbers.

Next: Registers — the sixteen slots where all of this arithmetic actually happens.