MODULE 03 / Days 21-30

Warps, reductions, atomics

All lessons
  1. 21
    What is a warp in CUDA (and why warp size is 32)What is a warp in CUDA, and why does a 48-thread block cost two warp slots?
    run logged
  2. 22
    Warp divergence and what it costsWhat is warp divergence in CUDA, and what does it cost?
    run logged
  3. 23
    __shfl_down_sync and the mask argumentWhat does the mask argument in __shfl_down_sync actually do?
    run logged
  4. 24
    Parallel reduction, versions 1 to 4How much is each step of the CUDA reduction ladder actually worth?
    run logged
  5. 25
    Optimizing a CUDA reduction to the memory ceilingHow do I make a CUDA reduction faster?
    run logged
  6. 26
    CUDA atomicAdd and what contention costsHow slow is atomicAdd in CUDA, and when is a reduction faster?
    run logged
  7. 27
    The CUDA memory model: fences, scopes and atomicsWhat does __threadfence() do in CUDA, and when do I need one?
    run logged
  8. 28
    CUDA cooperative groups and grid-wide syncHow do I synchronize every block in a CUDA grid?
    run logged
  9. 29
    CUDA histograms, privatized and measuredWhy is a CUDA histogram slow, and how does privatization fix it?
    run logged
  10. 30
    Arithmetic intensity and the memory-bound mindsetIs my CUDA kernel memory bound or compute bound, and how do I tell before I profile?
    run logged