MODULE 08 / Days 71-80
Tensor cores and modern hardware
- 71Mixed precision: FP16, BF16, TF32, FP8, FP4What is mixed precision in CUDA, and why accumulate in FP32?run logged
- 72Tensor cores with WMMAHow do I use tensor cores from CUDA C++ with the WMMA API?run logged
- 73mma.sync and ldmatrixHow do I use mma.sync and ldmatrix, and do they need an Ampere GPU?run logged
- 74Async copies and pipelinesWhat is cp.async, and how does cuda::pipeline use it?lesson
- 75Asynchronous barriersWhat is cuda::barrier and how is it different from __syncthreads?lesson
- 76The Tensor Memory Accelerator (TMA)How do I load a tile with TMA in CUDA?lesson
- 77Thread block clusters and distributed shared memoryWhat are thread block clusters and distributed shared memory in CUDA?lesson
- 78Which Blackwell do you have: wgmma, tcgen05 and what your card cannot doDoes an RTX 5090 support wgmma or tcgen05?lesson
- 79Tile programming with cuTileWhat is cuTile and how do I write a tile kernel?lesson
- 80Capstone 4: a tensor-core GEMMHow do I write a CUDA GEMM that uses tensor cores?run logged