MODULE 02 / Days 11-20
The memory hierarchy
- 11CUDA memory coalescing, measuredWhat is memory coalescing in CUDA, and when does it actually matter?run logged
- 12Why transposing a matrix is slowWhy is my CUDA matrix transpose so much slower than a copy?run logged
- 13CUDA shared memory and tiling, measured on a transposeHow do I use shared memory in CUDA, and how much does tiling actually buy?run logged
- 14What __syncthreads() guaranteesWhat does __syncthreads() actually guarantee, and why does my kernel work without it?run logged
- 15Shared memory bank conflicts explainedWhat is a shared memory bank conflict, and why does adding one column fix it?run logged
- 16CUDA tiled matrix multiplication, measuredHow does tiling speed up matrix multiplication in CUDA?run logged
- 17Registers, local memory and spillsHow do I tell whether my CUDA kernel is spilling registers?run logged
- 18CUDA constant memory, measured against globalIs CUDA constant memory faster than global memory?run logged
- 19Unified memory and what it costsIs cudaMallocManaged slower than cudaMalloc?run logged
- 20CUDA image convolution and edge detectionHow do I write a fast 2D convolution in CUDA?run logged