MODULE 08 / Days 71-80

Tensor cores and modern hardware

All lessons
  1. 71
    Mixed precision: FP16, BF16, TF32, FP8, FP4What is mixed precision in CUDA, and why accumulate in FP32?
    run logged
  2. 72
    Tensor cores with WMMAHow do I use tensor cores from CUDA C++ with the WMMA API?
    run logged
  3. 73
    mma.sync and ldmatrixHow do I use mma.sync and ldmatrix, and do they need an Ampere GPU?
    run logged
  4. 74
    Async copies and pipelinesWhat is cp.async, and how does cuda::pipeline use it?
    lesson
  5. 75
    Asynchronous barriersWhat is cuda::barrier and how is it different from __syncthreads?
    lesson
  6. 76
    The Tensor Memory Accelerator (TMA)How do I load a tile with TMA in CUDA?
    lesson
  7. 77
    Thread block clusters and distributed shared memoryWhat are thread block clusters and distributed shared memory in CUDA?
    lesson
  8. 78
    Which Blackwell do you have: wgmma, tcgen05 and what your card cannot doDoes an RTX 5090 support wgmma or tcgen05?
    lesson
  9. 79
    Tile programming with cuTileWhat is cuTile and how do I write a tile kernel?
    lesson
  10. 80
    Capstone 4: a tensor-core GEMMHow do I write a CUDA GEMM that uses tensor cores?
    run logged