Assembly Flavors
One Idea, Four Dialects
Take one trivial operation — add two numbers and return the result — and write it in each dialect. The differences you see here are the differences you will meet everywhere.
x86-64 (NASM, Intel syntax)
; long add(long a, long b) — System V ABI: rdi = a, rsi = b, result in rax
add_ints:
mov rax, rdi ; copy the first argument
add rax, rsi ; add the second: rax = a + b
ret
ARM64 (AArch64)
// long add(long a, long b) — AAPCS64: x0 = a, x1 = b, result in x0
add_ints:
add x0, x0, x1 // x0 = x0 + x1 (three-operand form)
ret // return via the link register (x30)
RISC-V (RV64)
# long add(long a, long b) — RISC-V calling convention: a0 = a, a1 = b
add_ints:
add a0, a0, a1 # a0 = a0 + a1
ret # return via the return-address register (ra)
WebAssembly (WAT)
(func $add_ints (param $a i64) (param $b i64) (result i64)
local.get $a ;; push a onto the stack
local.get $b ;; push b onto the stack
i64.add ;; pop both, push their sum
) ;; the remaining stack value is the result
Three observations, and they are the whole lesson:
- x86-64 uses two operands and destroys the destination, so you often need an extra
movfirst. - ARM64 and RISC-V use three operands (
add dst, src1, src2), which makes the intent clearer and avoids copying. - WebAssembly has no operands at all — it pushes and pops an implicit operand stack, which is why it needs no register names.
| Property | x86-64 | ARM64 | RISC-V | WASM |
|---|---|---|---|---|
| Design | CISC | RISC | RISC | Stack VM |
| Instruction size | 1–15 bytes | 4 bytes | 2–4 bytes | Variable bytecode |
| Operands in memory? | Yes | No | No | No (explicit load/store) |
| Registers (integer) | 16 | 31 | 32 | none (stack) |
| Flags register | dedicated | nzcv | none | none |
| Return address | on the stack | in lr | in ra | implicit call stack |
ARM64 in Detail
A Load/Store Architecture
ARM64 cannot use memory as an operand of an arithmetic instruction. Every value must be loaded into a register, operated on, and stored back. This single rule shapes all ARM assembly.
// sum of a 64-bit array — ARM64 (AArch64)
// void sum(const long *p, long n) -> x0 = p, x1 = n, result in x0
sum:
mov x2, #0 // x2 = total = 0
mov x3, #0 // x3 = index i = 0
loop:
cmp x3, x1 // compare i with n
b.ge done // branch if i >= n (b.)
ldr x4, [x0, x3, lsl #3] // load p[i]: base x0, index x3 scaled by 8
add x2, x2, x4 // total += p[i]
add x3, x3, #1 // i++
b loop
done:
mov x0, x2 // return total in x0
ret // return via the link register
Notice what replaced the x86 forms: ldr for the load, an explicit scale written as a shift (lsl #3 means ×8), and b.ge for a conditional branch. There is no mov rax, [rbx + rcx*8] equivalent — the load and the arithmetic are always separate instructions.
Conditional Instructions
ARM's flags (nzcv) are set only by instructions that explicitly end in s, and — unusually — many instructions can be executed conditionally without a branch at all:
cmp x0, #0 // compare
add x1, x1, #1 // this ALWAYS runs
cmp x0, #0
add.gt x1, x1, #1 // this runs ONLY if x0 > 0
csel x2, x3, x4, gt // x2 = (x0 > 0) ? x3 : x4 — branchless select
csel is the ARM equivalent of x86's cmov, and conditional execution is a genuine ARM strength: short if bodies often become a single conditional instruction with no branch to predict.
Calling into ARM
The procedure call standard passes the first eight arguments in x0–x7 and returns in x0. The return address lives in x30/lr, which means a non-leaf function must save it on the stack before calling anything else.
// leaf function: no stack needed
double_it:
lsl x0, x0, #1 // x0 = x0 << 1 (multiply by 2)
ret
// non-leaf function: must preserve lr
outer:
stp x29, x30, [sp, #-16]! // push frame pointer + link register
mov x29, sp
bl double_it // call (sets lr to the return address)
ldp x29, x30, [sp], #16 // restore them (post-increment)
ret
Run ARM64 code without an ARM machine using the CPUlator browser simulator, or in QEMU user-mode: qemu-aarch64 program. The demo programs in this track include ARM64 sources you can assemble with aarch64-linux-gnu-as.
RISC-V Essentials
RISC-V is the newest of the four dialects and the simplest: a small base instruction set (RV64I) plus standard extensions. It is also load/store, and its register names are deliberately unsurprising:
# Sum an array of 64-bit values — RISC-V RV64 (GNU as, AT&T-like syntax)
# long sum(const long *p, long n) -> a0 = p, a1 = n
sum:
li t0, 0 # t0 = total = 0
li t1, 0 # t1 = index i = 0
loop:
bge t1, a1, done # if i >= n, branch to done
slli t2, t1, 3 # t2 = i * 8 (shift left logical immediate)
add t2, a0, t2 # t2 = address of p[i]
ld t3, 0(t2) # t3 = *t2 (load doubleword)
add t0, t0, t3 # total += p[i]
addi t1, t1, 1 # i++
j loop
done:
mv a0, t0 # return value goes in a0
ret
Two RISC-V features are worth knowing. li ("load immediate") is a pseudo-instruction that the assembler expands into one or two real instructions depending on the constant. And mv is a pseudo-instruction for addi rd, rs, 0. Pseudo-instructions keep RISC-V source readable while the underlying ISA stays tiny.
WebAssembly Text Format
WebAssembly (WASM) is not a processor — it is a portable bytecode with a human-readable text form called WAT. It is included here because it is the dialect most web developers will encounter, and because it is genuinely different: WAT describes a stack machine with no named registers at all.
(module
;; memory: one page = 64 KiB, exported so JavaScript can read it
(memory (export "memory") 1)
;; data placed at offset 0, at instantiation time
(data (i32.const 0) "Hello from WASM!\n")
;; imported function from the host: (fd, ptr, len) -> written
(import "env" "write" (func $write (param i32 i32 i32) (result i32)))
;; exported function the host can call
(func (export "greet") (result i32)
i32.const 1 ;; push fd = 1 (stdout)
i32.const 0 ;; push ptr = 0 (start of the data above)
i32.const 16 ;; push len = 16 bytes
call $write ;; pop three arguments, push the result
)
)
Read it as a stack: i32.const 1 pushes a value, and call $write pops the three values it needs. There is no mov, no rax, and no syscall — those are replaced by operands on an implicit stack and by imported host functions. Compile it with wat2wasm (from the WABT toolkit) or run it directly in a browser's WebAssembly.instantiateStreaming.
Porting Between Dialects
You will rarely rewrite a whole program by hand, but you will read one dialect and think in another. The translation is mostly mechanical because the concepts are the same:
| Concept | x86-64 (NASM) | ARM64 | RISC-V |
|---|---|---|---|
| Move a constant | mov rax, 5 | mov x0, #5 | li a0, 5 |
| Add registers | add rax, rbx | add x0, x0, x1 | add a0, a0, a1 |
| Load from memory | mov rax, [rbx] | ldr x0, [x1] | ld a0, 0(a1) |
| Store to memory | mov [rbx], rax | str x0, [x1] | sd a0, 0(a1) |
| Call a function | call fn | bl fn | call fn |
| Return | ret | ret (uses lr) | ret (uses ra) |
| Branch if zero | jz / je | cbz | beqz |
| Unconditional branch | jmp | b | j |
Three habits make porting reliable: check the argument register names first (they differ per ABI), remember that RISC has no memory operands in arithmetic instructions, and re-read your conditional jumps, because the signed/unsigned distinction is expressed differently in every dialect.
Summary
- x86-64 is CISC: many instruction forms, memory operands allowed, variable-length encoding.
- ARM64 is RISC with a link register, fixed 32-bit instructions, and conditional instructions as well as conditional branches.
- RISC-V is a small, open RISC ISA whose base plus extensions make it easy to learn and to implement.
- WebAssembly is a stack machine bytecode with a text form (WAT) and no registers at all.
- The concepts — load/store, arithmetic, branches, calls — are identical everywhere; only the spelling and the ABI change.
Next: Lab Examples — runnable programs in x86-64 and ARM that put everything together.