GPU Computing

A GPU is a thousand-core co-processor with its own memory, driven by the CPU as host. While the multicore lab used the CPU's handful of powerful cores, this lab uses the GPU's thousands of small cores: the host + device model, kernel execution, the memory-transfer bottleneck, and the three frameworks — CUDA, OpenCL, and SYCL — that make C the host language of GPUs.

Why a GPU

A GPU trades single-thread speed for throughput: it runs thousands of lightweight threads at once, each doing a small slice of a huge, uniform workload. Problems with millions of independent elements — matrix math, image filters, machine-learning inference — map naturally onto that shape; problems with long dependency chains do not.

Multicore vs Many-Core

The CPU is multicore: 4–64 fast cores optimized for latency and complex control flow. The GPU is many-core: thousands of simple cores optimized for throughput and data-parallel math. A CPU core can branch, speculate, and cache; a GPU core wins when every thread executes the same instruction on different data — the SIMT (single instruction, multiple threads) model. If your hot loop has data-dependent branches everywhere, it belongs on the CPU; if it is a uniform transform, it belongs on the GPU.

The Host + Device Model

Every GPU program is a heterogeneous program: the CPU (host) runs ordinary C, owns the data, and the GPU (device) runs kernels. The two have separate memories — a key difference from the shared-memory tools of the multicore lab — so data must be copied across an explicit boundary before and after each kernel.

CPU host and GPU device with a PCIe link for transfers

Figure 1 — multicore CPU (host) plus GPU (device): kernels run on the GPU, data crosses the PCIe link as explicit copies.

Kernel Execution

A kernel is a function that runs on the device. When you launch it, the hardware instantiates it as a grid of blocks, each block a group of threads; every thread runs the same code on its own element index. The launch syntax kernel<<<blocks, threads>>>(...) is CUDA's way of saying "make blocks × threads copies". Inside the kernel, the thread asks the runtime who it is (blockIdx, blockDim, threadIdx) and indexes its private slice of the work.

Memory Transfers — The Bottleneck

Host and device memory are different address spaces: cudaMemcpy(..., cudaMemcpyHostToDevice) ships the input over the PCIe bus, the kernel runs, then cudaMemcpy(..., cudaMemcpyDeviceToHost) brings results back. That copy is the classic bottleneck — a fast kernel can be starved by a slow transfer. The rules: copy once, batch small transfers, overlap transfers with kernel execution (streams), and keep kernels long enough to amortize the cost.

Frameworks

Three frameworks dominate, and all three treat C/C++ as the host language. The choice is mostly about vendor support and portability:

CUDA — NVIDIA's Native Framework

CUDA extends C with kernel syntax (__global__ functions, the <<<grid, block>>> launch operator) and the cudaMemcpy API. It is the fastest path to NVIDIA hardware, the best tooling, and the most documentation — but it is NVIDIA-only. Kernels are compiled ahead of time by nvcc.

OpenCL — The Portable Standard

OpenCL is an open standard that runs on GPUs, CPUs, FPGAs, and DSPs from every vendor. The model is the same (host + kernels), but kernels are written as source strings compiled at runtime, and the host code is a verbose API (clCreateContext, clBuildProgram...). Portability comes at the cost of boilerplate and runtime compilation.

SYCL — Single-Source, C++-Based

SYCL keeps the host and kernel in one C++ source file, describing device work as command groups; an implementation (e.g. Intel's oneAPI DPC++) compiles for CPU, GPU, and FPGA from the same code. It is the modern choice when you want CUDA-like productivity without locking onto one vendor.

A First Kernel

Every GPU program is the same skeleton: allocate device memory, copy in, launch a kernel whose threads each own one element, copy the result back. The vector-add below is that skeleton in its smallest correct form — study the shape, not the arithmetic.

Vector Add (CUDA)

// kernel.cu — CUDA kernel + host code (C-like; built with nvcc).
#include <stdio.h>

#define N 1024

// __global__ marks a function that runs ON the device, in every thread
__global__ void add_kernel(int *c, const int *a, const int *b, int n) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;   // global thread id
    if (i < n) {
        c[i] = a[i] + b[i];   // each thread owns one element — no races
    }
}

int main(void) {
    int a[N], b[N], c[N];
    int *a_dev, *b_dev, *c_dev;
    for (int i = 0; i < N; i++) { a[i] = i; b[i] = i; }

    // 1) allocate device memory
    cudaMalloc(&a_dev, N * sizeof(int));
    cudaMalloc(&b_dev, N * sizeof(int));
    cudaMalloc(&c_dev, N * sizeof(int));

    // 2) copy inputs host -> device
    cudaMemcpy(a_dev, a, N * sizeof(int), cudaMemcpyHostToDevice);
    cudaMemcpy(b_dev, b, N * sizeof(int), cudaMemcpyHostToDevice);

    // 3) launch: 1 block x 1024 threads, all running add_kernel
    add_kernel<<<1, N>>>(c_dev, a_dev, b_dev, N);

    // 4) copy the result device -> host
    cudaMemcpy(c, c_dev, N * sizeof(int), cudaMemcpyDeviceToHost);

    // 5) verify one element and clean up
    printf("c[512] = %d (want %d)\n", c[512], 2 * 512);
    cudaFree(a_dev); cudaFree(b_dev); cudaFree(c_dev);
    return 0;
}

The guard if (i < n) matters: real grids rarely divide the work evenly, so a full-size launch must clamp the index instead of writing out of bounds. Real kernels scale the launch to e.g. <<<(N + 255) / 256, 256>>> — the ceiling division computes enough blocks to cover any N.

Building with nvcc

CUDA source compiles with NVIDIA's compiler driver, which splits host code (C) from device code (kernels):

nvcc -std=c++11 -O2 kernel.cu -o kernel   # note: host side is C++; kernels C-like
./kernel

OpenCL programs instead ship the kernel as a string (clBuildProgram compiles it when the program starts); SYCL/DPC++ again compiles ahead of time. Whichever framework you pick, the host language remains C-ish and the memory model is the one taught here.

When the GPU Wins

The GPU is not a faster CPU; it is a different shape of computer. It wins when:

  • The workload has massive data-level parallelism — millions of independent elements.
  • It is compute-bound math (matrix products, convolutions, simulations), not I/O or branchy logic.
  • Data transfer is amortized — the kernel runs long and infrequently compared with the copies.

It loses when the problem is small (copy overhead dominates), serial (a long dependency chain), or latency-critical (the driver overhead is microseconds). The engineering rule from the multicore lab still applies: profile first, GPU second. When you do commit, the profiling tools — NVIDIA nsight, ncu — will show you exactly where the memory-transfer tax is being paid.

Next lab: ABI & assembly — the machine-level contract beneath every function call.