COURSE / LESSONS

CUDA from day 0 to day 100

Open the learning path
00
1 lessons
  1. 0
    What you need to know before learning CUDAWhat do I need to know before I start learning CUDA?
    run logged
01

Days 1-10

First kernels

9 lessons
  1. 1
    Your first CUDA kernel (and why it prints nothing)Why does my CUDA hello world print nothing?
    run logged
  2. 2
    How a GPU differs from a CPUHow is a GPU different from a CPU?
    run logged
  3. 4
    CUDA grid, block and thread indexing explainedHow do blockIdx, blockDim and threadIdx map a thread to an array element?
    run logged
  4. 5
    CUDA vector addition, end to end with error checkingHow do I write a CUDA vector addition, and why does mine run without reporting an error?
    run logged
  5. 6
    How to check for errors in CUDAHow do I check for errors in CUDA, and why is the error reported at the wrong line?
    run logged
  6. 7
    Two-dimensional grids and image kernelsHow do I launch a 2D grid in CUDA, and does x map to the row or the column?
    run logged
  7. 8
    Bounds checks and CUDA grid-stride loopsHow do I write a CUDA kernel that is correct for any N?
    run logged
  8. 9
    Why your GPU code looks slower than your CPUWhy is my CUDA code slower than my CPU?
    run logged
  9. 10
    How to choose threads per block and blocks per gridHow do I choose the number of threads per block and blocks per grid for a CUDA kernel?
    run logged
02
10 lessons
  1. 11
    CUDA memory coalescing, measuredWhat is memory coalescing in CUDA, and when does it actually matter?
    run logged
  2. 12
    Why transposing a matrix is slowWhy is my CUDA matrix transpose so much slower than a copy?
    run logged
  3. 13
    CUDA shared memory and tiling, measured on a transposeHow do I use shared memory in CUDA, and how much does tiling actually buy?
    run logged
  4. 14
    What __syncthreads() guaranteesWhat does __syncthreads() actually guarantee, and why does my kernel work without it?
    run logged
  5. 15
    Shared memory bank conflicts explainedWhat is a shared memory bank conflict, and why does adding one column fix it?
    run logged
  6. 16
    CUDA tiled matrix multiplication, measuredHow does tiling speed up matrix multiplication in CUDA?
    run logged
  7. 17
    Registers, local memory and spillsHow do I tell whether my CUDA kernel is spilling registers?
    run logged
  8. 18
    CUDA constant memory, measured against globalIs CUDA constant memory faster than global memory?
    run logged
  9. 19
    Unified memory and what it costsIs cudaMallocManaged slower than cudaMalloc?
    run logged
  10. 20
    CUDA image convolution and edge detectionHow do I write a fast 2D convolution in CUDA?
    run logged
03
10 lessons
  1. 21
    What is a warp in CUDA (and why warp size is 32)What is a warp in CUDA, and why does a 48-thread block cost two warp slots?
    run logged
  2. 22
    Warp divergence and what it costsWhat is warp divergence in CUDA, and what does it cost?
    run logged
  3. 23
    __shfl_down_sync and the mask argumentWhat does the mask argument in __shfl_down_sync actually do?
    run logged
  4. 24
    Parallel reduction, versions 1 to 4How much is each step of the CUDA reduction ladder actually worth?
    run logged
  5. 25
    Optimizing a CUDA reduction to the memory ceilingHow do I make a CUDA reduction faster?
    run logged
  6. 26
    CUDA atomicAdd and what contention costsHow slow is atomicAdd in CUDA, and when is a reduction faster?
    run logged
  7. 27
    The CUDA memory model: fences, scopes and atomicsWhat does __threadfence() do in CUDA, and when do I need one?
    run logged
  8. 28
    CUDA cooperative groups and grid-wide syncHow do I synchronize every block in a CUDA grid?
    run logged
  9. 29
    CUDA histograms, privatized and measuredWhy is a CUDA histogram slow, and how does privatization fix it?
    run logged
  10. 30
    Arithmetic intensity and the memory-bound mindsetIs my CUDA kernel memory bound or compute bound, and how do I tell before I profile?
    run logged
04
10 lessons
  1. 31
    CUDA prefix sum, the Hillis-Steele scanHow do I write a prefix sum in CUDA, and why does it need two shared arrays?
    run logged
  2. 32
    The CUDA scan algorithm: the tree and the three kernelsHow does a work-efficient CUDA scan work, and why three kernels?
    run logged
  3. 33
    CUDA stream compaction with a prefix sumHow do I remove elements from an array on a GPU?
    run logged
  4. 34
    CUDA stencils and the heat equationHow do I write a 2D stencil in CUDA, and why does my heat simulation blow up?
    run logged
  5. 35
    CUDA radix sort, built from the scan you already wroteHow does a radix sort work on a GPU, and why is it the sort GPUs use?
    run logged
  6. 36
    CUDA merge path: merging sorted arrays in parallelHow do you merge two sorted arrays in parallel on a GPU?
    run logged
  7. 37
    Sparse matrices in CUDA: COO, CSR and SpMVWhat is CSR, and why is my CUDA SpMV slow on some matrices?
    run logged
  8. 38
    Graph traversal on a GPUHow do you run a breadth-first search on a GPU, and why is one thread per vertex slow?
    run logged
  9. 39
    When to stop hand-writing: Thrust, CUB and libcu++Should I use Thrust and CUB or write the CUDA kernel myself?
    run logged
  10. 40
    PageRank on a real graph in CUDAHow do I implement PageRank in CUDA?
    run logged
05
10 lessons
  1. 41
    Reading an Nsight Systems timelineHow do I read an Nsight Systems timeline?
    run logged
  2. 42
    Reading an Nsight Compute reportHow do I read an Nsight Compute report?
    run logged
  3. 43
    Optimizing CUDA matmul, steps 1 to 3How do I optimize a CUDA matrix multiplication kernel?
    run logged
  4. 44
    Optimizing CUDA matmul, steps 4 to 6How do I get a CUDA matmul close to cuBLAS speed?
    run logged
  5. 45
    Occupancy is not the goalDoes higher occupancy make a CUDA kernel faster?
    run logged
  6. 46
    Reading PTX and SASS, and which one runsWhat is the difference between PTX and SASS in CUDA?
    run logged
  7. 47
    Fast math, FMA and precision, one flag at a timeWhat does -use_fast_math actually do in CUDA, and what does it cost?
    run logged
  8. 48
    CUDA kernel fusion and what a launch really costsWhen does fusing CUDA kernels actually make them faster?
    run logged
  9. 49
    Building a roofline for your GPUHow do I build a roofline model for my own GPU?
    run logged
  10. 50
    The CUDA performance checklist, in diagnostic orderMy CUDA kernel is slow. What do I check first?
    run logged
06

Days 51-60

Concurrency

10 lessons
  1. 51
    CUDA streams and overlapHow do CUDA streams overlap two kernels?
    run logged
  2. 52
    Events and cross-stream dependenciesHow do I make one CUDA stream wait on another?
    run logged
  3. 53
    Pinned memory and async copiesWhy is cudaMemcpyAsync not asynchronous?
    run logged
  4. 54
    Double bufferingHow does double buffering overlap CUDA copies and kernels?
    run logged
  5. 55
    Stream-ordered allocation and memory poolsWhat does cudaMallocAsync actually do?
    run logged
  6. 56
    CUDA graphsWhen do CUDA graphs actually make a program faster?
    run logged
  7. 57
    Graph updates and conditional nodesHow do I change a CUDA graph's parameters, and can a graph loop until convergence without the host?
    run logged
  8. 58
    Programmatic dependent launchWhat is programmatic dependent launch in CUDA?
    lesson
  9. 59
    Host threads and the GPUCan multiple CPU threads use CUDA at the same time?
    run logged
  10. 60
    Capstone 3: a real-time frame pipelineHow do I build a real-time video pipeline in CUDA?
    run logged
07
9 lessons
  1. 61
    Finding memory bugs with compute-sanitizerHow do I find out-of-bounds accesses and memory leaks in a CUDA kernel?
    run logged
  2. 62
    Finding races with racecheck and synccheckHow do I find a data race in a CUDA kernel?
    run logged
  3. 63
    Debugging a kernel with cuda-gdbHow do I debug a CUDA kernel with cuda-gdb?
    run logged
  4. 64
    printf, assert and when printf liesWhy does printf in my CUDA kernel print nothing?
    run logged
  5. 65
    Deadlocks and hangsWhy does my CUDA kernel hang, and how do I find out where?
    run logged
  6. 66
    Testing GPU kernelsHow do I test CUDA kernels?
    run logged
  7. 68
    Floating point and reproducibilityWhy does my CUDA reduction give a slightly different result every run?
    run logged
  8. 69
    One CUDA binary across GPU generationsHow do I build one CUDA binary that runs on every GPU generation?
    run logged
  9. 70
    Checkpoint: the CUDA bug catalogWhich tool catches which CUDA bug, and which bugs does no tool see?
    run logged
08
10 lessons
  1. 71
    Mixed precision: FP16, BF16, TF32, FP8, FP4What is mixed precision in CUDA, and why accumulate in FP32?
    run logged
  2. 72
    Tensor cores with WMMAHow do I use tensor cores from CUDA C++ with the WMMA API?
    run logged
  3. 73
    mma.sync and ldmatrixHow do I use mma.sync and ldmatrix, and do they need an Ampere GPU?
    run logged
  4. 74
    Async copies and pipelinesWhat is cp.async, and how does cuda::pipeline use it?
    lesson
  5. 75
    Asynchronous barriersWhat is cuda::barrier and how is it different from __syncthreads?
    lesson
  6. 76
    The Tensor Memory Accelerator (TMA)How do I load a tile with TMA in CUDA?
    lesson
  7. 77
    Thread block clusters and distributed shared memoryWhat are thread block clusters and distributed shared memory in CUDA?
    lesson
  8. 78
    Which Blackwell do you have: wgmma, tcgen05 and what your card cannot doDoes an RTX 5090 support wgmma or tcgen05?
    lesson
  9. 79
    Tile programming with cuTileWhat is cuTile and how do I write a tile kernel?
    lesson
  10. 80
    Capstone 4: a tensor-core GEMMHow do I write a CUDA GEMM that uses tensor cores?
    run logged
09
9 lessons
  1. 81
    cuBLAS and cuBLASLtHow do I call cuBLAS from C++ when my matrices are row major?
    run logged
  2. 82
    cuFFT, cuRAND, cuSPARSE and cuSOLVERWhich CUDA math library replaces the kernel I was about to write?
    run logged
  3. 83
    cuDNN with the graph APIHow do I call cuDNN's graph API for a Conv2D forward?
    run logged
  4. 84
    CUTLASS and CuTe layoutsWhat is CuTe, and does CUTLASS need an Ampere GPU?
    run logged
  5. 85
    CUDA from Python: cuda.core, CuPy, Numba and PyCUDAHow do I launch a CUDA kernel from Python?
    run logged
  6. 86
    Writing kernels in TritonHow do I write a CUDA kernel in Triton, and will it run on my GPU?
    lesson
  7. 87
    Custom PyTorch operatorsHow do I call my own CUDA kernel from PyTorch with a backward pass?
    run logged
  8. 88
    Profiling a PyTorch model down to the kernelHow do I profile a PyTorch model and find the slowest CUDA kernel?
    run logged
  9. 90
    Checkpoint: library or kernel?When should I write a custom CUDA kernel instead of calling a library?
    run logged
10
10 lessons
  1. 91
    Using more than one GPUHow do I use two GPUs in one CUDA program?
    lesson
  2. 92
    NCCL and collective operationsHow does an NCCL all-reduce work, and why does adding GPUs not make it worse?
    lesson
  3. 93
    Unified memory, second passWhat does cudaMemAdvise actually do, and when does it beat a prefetch?
    run logged
  4. 94
    Sharing a GPU: green contexts, MPS and MIGHow do I run two CUDA workloads on one GPU without them fighting?
    run logged
  5. 95
    LLM kernels 1: softmax, layer norm, RMS normHow do you write a fast softmax, layer norm and RMS norm kernel in CUDA?
    run logged
  6. 96
    LLM kernels 2: RoPE, GELU and elementwise fusionHow do you write a RoPE and a GELU kernel in CUDA, and what does fusing them save?
    run logged
  7. 97
    LLM kernels 3: attention and online softmaxHow does FlashAttention compute attention without storing the N by N matrix?
    run logged
  8. 98
    Quantization and sparsityHow do you write an INT8 matmul in CUDA, and where does the scale go?
    run logged
  9. 99
    Assembling a forward and a backward passHow do you write and check a neural-network backward pass in CUDA?
    run logged
  10. 100
    Capstone 5: your own llm.c-liteHow do I train a neural network with CUDA kernels I wrote myself?
    run logged