MODULE 10 / Days 91-100
Scale and the final capstone
- 91Using more than one GPUHow do I use two GPUs in one CUDA program?lesson
- 92NCCL and collective operationsHow does an NCCL all-reduce work, and why does adding GPUs not make it worse?lesson
- 93Unified memory, second passWhat does cudaMemAdvise actually do, and when does it beat a prefetch?run logged
- 94Sharing a GPU: green contexts, MPS and MIGHow do I run two CUDA workloads on one GPU without them fighting?run logged
- 95LLM kernels 1: softmax, layer norm, RMS normHow do you write a fast softmax, layer norm and RMS norm kernel in CUDA?run logged
- 96LLM kernels 2: RoPE, GELU and elementwise fusionHow do you write a RoPE and a GELU kernel in CUDA, and what does fusing them save?run logged
- 97LLM kernels 3: attention and online softmaxHow does FlashAttention compute attention without storing the N by N matrix?run logged
- 98Quantization and sparsityHow do you write an INT8 matmul in CUDA, and where does the scale go?run logged
- 99Assembling a forward and a backward passHow do you write and check a neural-network backward pass in CUDA?run logged
- 100Capstone 5: your own llm.c-liteHow do I train a neural network with CUDA kernels I wrote myself?run logged