MODULE 03 / Days 21-30
Warps, reductions, atomics
- 21What is a warp in CUDA (and why warp size is 32)What is a warp in CUDA, and why does a 48-thread block cost two warp slots?run logged
- 22Warp divergence and what it costsWhat is warp divergence in CUDA, and what does it cost?run logged
- 23__shfl_down_sync and the mask argumentWhat does the mask argument in __shfl_down_sync actually do?run logged
- 24Parallel reduction, versions 1 to 4How much is each step of the CUDA reduction ladder actually worth?run logged
- 25Optimizing a CUDA reduction to the memory ceilingHow do I make a CUDA reduction faster?run logged
- 26CUDA atomicAdd and what contention costsHow slow is atomicAdd in CUDA, and when is a reduction faster?run logged
- 27The CUDA memory model: fences, scopes and atomicsWhat does __threadfence() do in CUDA, and when do I need one?run logged
- 28CUDA cooperative groups and grid-wide syncHow do I synchronize every block in a CUDA grid?run logged
- 29CUDA histograms, privatized and measuredWhy is a CUDA histogram slow, and how does privatization fix it?run logged
- 30Arithmetic intensity and the memory-bound mindsetIs my CUDA kernel memory bound or compute bound, and how do I tell before I profile?run logged