Study Projects
Project 4 — ARM64 String Reverse
Goal: reverse a string in place on ARM64, then print it with the write syscall. The algorithm is identical to x86; only the spelling of the loads, stores, and branches changes.
Combines: the ARM64 material from lesson 14, plus every ARM demo in the lab. Start from demo 20.
// reverse(buffer, len) — in place
// x0 = buffer address, x1 = length in bytes
reverse:
mov x2, x0 // x2 = left = buffer
add x3, x0, x1 // x3 = right = buffer + len
sub x3, x3, #1 // point at the LAST byte
swap_loop:
cmp x2, x3 // have the pointers met or crossed?
b.ge done
ldrb w4, [x2] // w4 = *left (byte load: ldrb)
ldrb w5, [x3] // w5 = *right
strb w5, [x2] // *left = w5 (byte store: strb)
strb w4, [x3] // *right = w4
add x2, x2, #1 // left++
sub x3, x3, #1 // right--
b swap_loop
done:
ret
Watch out for two things. First, ARM64 requires byte-sized loads and stores here — ldrb/strb, not ldr/str, which would move eight bytes. Second, this routine touches no memory beyond the buffer, so it is a leaf function and needs no stack frame at all.
rsi, a right index in rdi, and use movzx to read each byte before storing it with mov byte [ptr], reg. Writing the same routine twice, in two dialects, is the single most effective exercise in this track.
How to Work Through a Project
- Make it run first. Get the starting demo working unmodified before adding anything.
- Change one thing. Add one instruction, rebuild, run. Confirm the output changed as predicted.
- Instrument it. When a value is wrong, print it. In assembly, printing is debugging.
- Read the disassembly.
objdump -dshows what you actually wrote — useful when the meaning of a line surprises you. - Write the comment before the code. Decide the algorithm in English, then translate line by line.
Summary
- Projects combine lessons; each one here extends a demo you have already run.
- Hex dump teaches syscalls, loops, and table lookup.
- Bubble sort teaches scaled indexing and the signed/unsigned branch choice.
- The C bridge teaches the ABI, stack alignment, and libc interoperability.
- The ARM64 string reverse teaches the same algorithm in a load/store dialect — and proves the concepts transfer.
Next: References & Assemblers — every tool, manual, and course you need to keep going.
Project 1 — Hex Dump
Goal: read bytes from standard input and print them as hexadecimal, sixteen per line, with the printable ASCII shown alongside — the output you see from xxd or hexdump -C.
Combines: the read syscall (lesson 12), a byte loop, and a hex-digit conversion. Start from
demo 12 and the number printer in
demo 13.
The core of the work is converting one byte into two hex characters. Because a nibble is exactly four bits, a 16-entry lookup table removes all arithmetic:
section .rodata
HEX db "0123456789ABCDEF" ; index 0..15 -> the matching character
section .text
; al holds the byte to display
mov rbx, HEX ; base of the table
movzx ecx, al ; ecx = the byte, zero-extended
shr ecx, 4 ; high nibble (0..15)
mov dl, [rbx + rcx] ; look up the first hex digit
; ... output dl ...
movzx ecx, al
and ecx, 0x0F ; low nibble (0..15)
mov dl, [rbx + rcx] ; look up the second hex digit
; ... output dl ...
Stretch goal: buffer a whole line of output in .bss and issue a single write per line instead of one per character. Measure both versions with time on a 1 MB input — the difference is the cost of a system call, and it is enormous.
Project 2 — Bubble Sort
Goal: sort an array of signed 64-bit integers in place, then print it. Implement the swap, the outer pass, and the early-exit "already sorted" flag.
Combines: scaled indexing (lesson 6), the compare-and-jump family (lesson 9), and the number printer. Start from demo 14.
; for (j = 0; j < n-1; j++)
; if (a[j] > a[j+1]) swap them
;
; a is in rbx, n in r12, j in rcx
.inner:
mov rax, [rbx + rcx*8] ; a[j]
mov rdx, [rbx + rcx*8 + 8] ; a[j+1] (next element: +8 bytes)
cmp rax, rdx
jle .no_swap ; SIGNED compare — the array is signed!
mov [rbx + rcx*8], rdx ; swap
mov [rbx + rcx*8 + 8], rax
.no_swap:
inc rcx
The lesson in this project: change jle to jbe and watch negative numbers sort in the wrong place. Signedness is not academic — it decides whether your array is ordered correctly.
Project 3 — Calling C from Assembly
Goal: write main in assembly, call printf and malloc from the C library, and free the memory before exiting. This is the most realistic use of assembly in a normal project.
Combines: the ABI (lessons 4 and 10), stack alignment, and libc I/O. The pattern is the second half of
demo 9 after you convert it to use printf.
extern printf, malloc, free
; printf("%ld\n", value)
lea rdi, [rel fmt] ; argument 1: the format string
mov rsi, [value] ; argument 2: the number (note: NOT rax!)
xor eax, eax ; 0 vector registers used
call printf
; p = malloc(1024)
mov rdi, 1024
call malloc
mov [p], rax ; save the pointer
; free(p)
mov rdi, [p]
call free
Stretch goal: implement the same logic with raw syscalls (mmap instead of malloc, the number printer instead of printf) and compare the executable sizes with ls -l. The freestanding version is typically a few kilobytes; the libc version is tens of kilobytes and dynamically linked.