Assembly Flavors

Every dialect answers the same questions — where are the arguments, how do I load memory, how do I branch — with different spellings. Seeing one program in four dialects is the fastest way to stop being afraid of the others.

One Idea, Four Dialects

Take one trivial operation — add two numbers and return the result — and write it in each dialect. The differences you see here are the differences you will meet everywhere.

x86-64 (NASM, Intel syntax)

; long add(long a, long b)   — System V ABI: rdi = a, rsi = b, result in rax
add_ints:
    mov     rax, rdi        ; copy the first argument
    add     rax, rsi        ; add the second: rax = a + b
    ret

ARM64 (AArch64)

// long add(long a, long b)   — AAPCS64: x0 = a, x1 = b, result in x0
add_ints:
    add     x0, x0, x1      // x0 = x0 + x1  (three-operand form)
    ret                     // return via the link register (x30)

RISC-V (RV64)

# long add(long a, long b)   — RISC-V calling convention: a0 = a, a1 = b
add_ints:
    add     a0, a0, a1      # a0 = a0 + a1
    ret                     # return via the return-address register (ra)

WebAssembly (WAT)

(func $add_ints (param $a i64) (param $b i64) (result i64)
  local.get $a              ;; push a onto the stack
  local.get $b              ;; push b onto the stack
  i64.add                   ;; pop both, push their sum
)                           ;; the remaining stack value is the result

Three observations, and they are the whole lesson:

  • x86-64 uses two operands and destroys the destination, so you often need an extra mov first.
  • ARM64 and RISC-V use three operands (add dst, src1, src2), which makes the intent clearer and avoids copying.
  • WebAssembly has no operands at all — it pushes and pops an implicit operand stack, which is why it needs no register names.
Propertyx86-64ARM64RISC-VWASM
DesignCISCRISCRISCStack VM
Instruction size1–15 bytes4 bytes2–4 bytesVariable bytecode
Operands in memory?YesNoNoNo (explicit load/store)
Registers (integer)163132none (stack)
Flags registerdedicatednzcvnonenone
Return addresson the stackin lrin raimplicit call stack

ARM64 in Detail

A Load/Store Architecture

ARM64 cannot use memory as an operand of an arithmetic instruction. Every value must be loaded into a register, operated on, and stored back. This single rule shapes all ARM assembly.

// sum of a 64-bit array — ARM64 (AArch64)
// void sum(const long *p, long n)   -> x0 = p, x1 = n, result in x0
sum:
    mov     x2, #0              // x2 = total = 0
    mov     x3, #0              // x3 = index i = 0

loop:
    cmp     x3, x1              // compare i with n
    b.ge    done                // branch if i >= n  (b.)

    ldr     x4, [x0, x3, lsl #3]  // load p[i]: base x0, index x3 scaled by 8
    add     x2, x2, x4            // total += p[i]

    add     x3, x3, #1            // i++
    b       loop

done:
    mov     x0, x2              // return total in x0
    ret                         // return via the link register

Notice what replaced the x86 forms: ldr for the load, an explicit scale written as a shift (lsl #3 means ×8), and b.ge for a conditional branch. There is no mov rax, [rbx + rcx*8] equivalent — the load and the arithmetic are always separate instructions.

Conditional Instructions

ARM's flags (nzcv) are set only by instructions that explicitly end in s, and — unusually — many instructions can be executed conditionally without a branch at all:

    cmp     x0, #0              // compare
    add     x1, x1, #1          // this ALWAYS runs

    cmp     x0, #0
    add.gt  x1, x1, #1          // this runs ONLY if x0 > 0

    csel    x2, x3, x4, gt      // x2 = (x0 > 0) ? x3 : x4  — branchless select

csel is the ARM equivalent of x86's cmov, and conditional execution is a genuine ARM strength: short if bodies often become a single conditional instruction with no branch to predict.

Calling into ARM

The procedure call standard passes the first eight arguments in x0–x7 and returns in x0. The return address lives in x30/lr, which means a non-leaf function must save it on the stack before calling anything else.

// leaf function: no stack needed
double_it:
    lsl     x0, x0, #1          // x0 = x0 << 1  (multiply by 2)
    ret

// non-leaf function: must preserve lr
outer:
    stp     x29, x30, [sp, #-16]!   // push frame pointer + link register
    mov     x29, sp

    bl      double_it               // call (sets lr to the return address)

    ldp     x29, x30, [sp], #16     // restore them (post-increment)
    ret

Run ARM64 code without an ARM machine using the CPUlator browser simulator, or in QEMU user-mode: qemu-aarch64 program. The demo programs in this track include ARM64 sources you can assemble with aarch64-linux-gnu-as.

RISC-V Essentials

RISC-V is the newest of the four dialects and the simplest: a small base instruction set (RV64I) plus standard extensions. It is also load/store, and its register names are deliberately unsurprising:

# Sum an array of 64-bit values — RISC-V RV64 (GNU as, AT&T-like syntax)
# long sum(const long *p, long n)   -> a0 = p, a1 = n
sum:
    li      t0, 0               # t0 = total = 0
    li      t1, 0               # t1 = index i = 0
loop:
    bge     t1, a1, done        # if i >= n, branch to done

    slli    t2, t1, 3           # t2 = i * 8   (shift left logical immediate)
    add     t2, a0, t2          # t2 = address of p[i]
    ld      t3, 0(t2)           # t3 = *t2  (load doubleword)

    add     t0, t0, t3          # total += p[i]
    addi    t1, t1, 1           # i++
    j       loop
done:
    mv      a0, t0              # return value goes in a0
    ret

Two RISC-V features are worth knowing. li ("load immediate") is a pseudo-instruction that the assembler expands into one or two real instructions depending on the constant. And mv is a pseudo-instruction for addi rd, rs, 0. Pseudo-instructions keep RISC-V source readable while the underlying ISA stays tiny.

WebAssembly Text Format

WebAssembly (WASM) is not a processor — it is a portable bytecode with a human-readable text form called WAT. It is included here because it is the dialect most web developers will encounter, and because it is genuinely different: WAT describes a stack machine with no named registers at all.

(module
  ;; memory: one page = 64 KiB, exported so JavaScript can read it
  (memory (export "memory") 1)

  ;; data placed at offset 0, at instantiation time
  (data (i32.const 0) "Hello from WASM!\n")

  ;; imported function from the host: (fd, ptr, len) -> written
  (import "env" "write" (func $write (param i32 i32 i32) (result i32)))

  ;; exported function the host can call
  (func (export "greet") (result i32)
    i32.const 1        ;; push fd    = 1  (stdout)
    i32.const 0        ;; push ptr   = 0  (start of the data above)
    i32.const 16       ;; push len   = 16 bytes
    call $write        ;; pop three arguments, push the result
  )
)

Read it as a stack: i32.const 1 pushes a value, and call $write pops the three values it needs. There is no mov, no rax, and no syscall — those are replaced by operands on an implicit stack and by imported host functions. Compile it with wat2wasm (from the WABT toolkit) or run it directly in a browser's WebAssembly.instantiateStreaming.

Porting Between Dialects

You will rarely rewrite a whole program by hand, but you will read one dialect and think in another. The translation is mostly mechanical because the concepts are the same:

Conceptx86-64 (NASM)ARM64RISC-V
Move a constantmov rax, 5mov x0, #5li a0, 5
Add registersadd rax, rbxadd x0, x0, x1add a0, a0, a1
Load from memorymov rax, [rbx]ldr x0, [x1]ld a0, 0(a1)
Store to memorymov [rbx], raxstr x0, [x1]sd a0, 0(a1)
Call a functioncall fnbl fncall fn
Returnretret (uses lr)ret (uses ra)
Branch if zerojz / jecbzbeqz
Unconditional branchjmpbj

Three habits make porting reliable: check the argument register names first (they differ per ABI), remember that RISC has no memory operands in arithmetic instructions, and re-read your conditional jumps, because the signed/unsigned distinction is expressed differently in every dialect.

Summary

  • x86-64 is CISC: many instruction forms, memory operands allowed, variable-length encoding.
  • ARM64 is RISC with a link register, fixed 32-bit instructions, and conditional instructions as well as conditional branches.
  • RISC-V is a small, open RISC ISA whose base plus extensions make it easy to learn and to implement.
  • WebAssembly is a stack machine bytecode with a text form (WAT) and no registers at all.
  • The concepts — load/store, arithmetic, branches, calls — are identical everywhere; only the spelling and the ABI change.

Next: Lab Examples — runnable programs in x86-64 and ARM that put everything together.