Study Projects

Lessons teach ideas; projects build skill. These four projects each combine several lessons, and each one starts from a demo program you can already run — extend it, break it, and measure what changes.

Project 4 — ARM64 String Reverse

Goal: reverse a string in place on ARM64, then print it with the write syscall. The algorithm is identical to x86; only the spelling of the loads, stores, and branches changes.

Combines: the ARM64 material from lesson 14, plus every ARM demo in the lab. Start from demo 20.

// reverse(buffer, len)  — in place
// x0 = buffer address, x1 = length in bytes
reverse:
    mov     x2, x0              // x2 = left  = buffer
    add     x3, x0, x1          // x3 = right = buffer + len
    sub     x3, x3, #1          //           point at the LAST byte

swap_loop:
    cmp     x2, x3              // have the pointers met or crossed?
    b.ge    done

    ldrb    w4, [x2]            // w4 = *left    (byte load: ldrb)
    ldrb    w5, [x3]            // w5 = *right
    strb    w5, [x2]            // *left  = w5   (byte store: strb)
    strb    w4, [x3]            // *right = w4

    add     x2, x2, #1          // left++
    sub     x3, x3, #1          // right--
    b       swap_loop

done:
    ret

Watch out for two things. First, ARM64 requires byte-sized loads and stores here — ldrb/strb, not ldr/str, which would move eight bytes. Second, this routine touches no memory beyond the buffer, so it is a leaf function and needs no stack frame at all.

Snapshot: reimplement this on x86-64 as a final comparison — keep a left index in rsi, a right index in rdi, and use movzx to read each byte before storing it with mov byte [ptr], reg. Writing the same routine twice, in two dialects, is the single most effective exercise in this track.

How to Work Through a Project

  1. Make it run first. Get the starting demo working unmodified before adding anything.
  2. Change one thing. Add one instruction, rebuild, run. Confirm the output changed as predicted.
  3. Instrument it. When a value is wrong, print it. In assembly, printing is debugging.
  4. Read the disassembly. objdump -d shows what you actually wrote — useful when the meaning of a line surprises you.
  5. Write the comment before the code. Decide the algorithm in English, then translate line by line.

Summary

  • Projects combine lessons; each one here extends a demo you have already run.
  • Hex dump teaches syscalls, loops, and table lookup.
  • Bubble sort teaches scaled indexing and the signed/unsigned branch choice.
  • The C bridge teaches the ABI, stack alignment, and libc interoperability.
  • The ARM64 string reverse teaches the same algorithm in a load/store dialect — and proves the concepts transfer.

Next: References & Assemblers — every tool, manual, and course you need to keep going.

Project 1 — Hex Dump

Goal: read bytes from standard input and print them as hexadecimal, sixteen per line, with the printable ASCII shown alongside — the output you see from xxd or hexdump -C.

Combines: the read syscall (lesson 12), a byte loop, and a hex-digit conversion. Start from demo 12 and the number printer in demo 13.

The core of the work is converting one byte into two hex characters. Because a nibble is exactly four bits, a 16-entry lookup table removes all arithmetic:

section .rodata
    HEX  db "0123456789ABCDEF"       ; index 0..15 -> the matching character

section .text
    ; al holds the byte to display
    mov  rbx, HEX                    ; base of the table
    movzx ecx, al                    ; ecx = the byte, zero-extended
    shr  ecx, 4                      ; high nibble  (0..15)
    mov  dl, [rbx + rcx]             ; look up the first hex digit
    ; ... output dl ...
    movzx ecx, al
    and  ecx, 0x0F                   ; low nibble   (0..15)
    mov  dl, [rbx + rcx]             ; look up the second hex digit
    ; ... output dl ...

Stretch goal: buffer a whole line of output in .bss and issue a single write per line instead of one per character. Measure both versions with time on a 1 MB input — the difference is the cost of a system call, and it is enormous.

Project 2 — Bubble Sort

Goal: sort an array of signed 64-bit integers in place, then print it. Implement the swap, the outer pass, and the early-exit "already sorted" flag.

Combines: scaled indexing (lesson 6), the compare-and-jump family (lesson 9), and the number printer. Start from demo 14.

    ; for (j = 0; j < n-1; j++)
    ;     if (a[j] > a[j+1]) swap them
    ;
    ; a is in rbx, n in r12, j in rcx
.inner:
    mov  rax, [rbx + rcx*8]          ; a[j]
    mov  rdx, [rbx + rcx*8 + 8]      ; a[j+1]  (next element: +8 bytes)
    cmp  rax, rdx
    jle  .no_swap                    ; SIGNED compare — the array is signed!
    mov  [rbx + rcx*8], rdx          ; swap
    mov  [rbx + rcx*8 + 8], rax
.no_swap:
    inc  rcx

The lesson in this project: change jle to jbe and watch negative numbers sort in the wrong place. Signedness is not academic — it decides whether your array is ordered correctly.

Project 3 — Calling C from Assembly

Goal: write main in assembly, call printf and malloc from the C library, and free the memory before exiting. This is the most realistic use of assembly in a normal project.

Combines: the ABI (lessons 4 and 10), stack alignment, and libc I/O. The pattern is the second half of demo 9 after you convert it to use printf.

    extern printf, malloc, free

    ; printf("%ld\n", value)
    lea  rdi, [rel fmt]           ; argument 1: the format string
    mov  rsi, [value]             ; argument 2: the number (note: NOT rax!)
    xor  eax, eax                 ; 0 vector registers used
    call printf

    ; p = malloc(1024)
    mov  rdi, 1024
    call malloc
    mov  [p], rax                 ; save the pointer

    ; free(p)
    mov  rdi, [p]
    call free

Stretch goal: implement the same logic with raw syscalls (mmap instead of malloc, the number printer instead of printf) and compare the executable sizes with ls -l. The freestanding version is typically a few kilobytes; the libc version is tens of kilobytes and dynamically linked.