MODULE 10 / Days 91-100

Scale and the final capstone

All lessons
  1. 91
    Using more than one GPUHow do I use two GPUs in one CUDA program?
    lesson
  2. 92
    NCCL and collective operationsHow does an NCCL all-reduce work, and why does adding GPUs not make it worse?
    lesson
  3. 93
    Unified memory, second passWhat does cudaMemAdvise actually do, and when does it beat a prefetch?
    run logged
  4. 94
    Sharing a GPU: green contexts, MPS and MIGHow do I run two CUDA workloads on one GPU without them fighting?
    run logged
  5. 95
    LLM kernels 1: softmax, layer norm, RMS normHow do you write a fast softmax, layer norm and RMS norm kernel in CUDA?
    run logged
  6. 96
    LLM kernels 2: RoPE, GELU and elementwise fusionHow do you write a RoPE and a GELU kernel in CUDA, and what does fusing them save?
    run logged
  7. 97
    LLM kernels 3: attention and online softmaxHow does FlashAttention compute attention without storing the N by N matrix?
    run logged
  8. 98
    Quantization and sparsityHow do you write an INT8 matmul in CUDA, and where does the scale go?
    run logged
  9. 99
    Assembling a forward and a backward passHow do you write and check a neural-network backward pass in CUDA?
    run logged
  10. 100
    Capstone 5: your own llm.c-liteHow do I train a neural network with CUDA kernels I wrote myself?
    run logged