MODULE 02 / Days 11-20

The memory hierarchy

All lessons
  1. 11
    CUDA memory coalescing, measuredWhat is memory coalescing in CUDA, and when does it actually matter?
    run logged
  2. 12
    Why transposing a matrix is slowWhy is my CUDA matrix transpose so much slower than a copy?
    run logged
  3. 13
    CUDA shared memory and tiling, measured on a transposeHow do I use shared memory in CUDA, and how much does tiling actually buy?
    run logged
  4. 14
    What __syncthreads() guaranteesWhat does __syncthreads() actually guarantee, and why does my kernel work without it?
    run logged
  5. 15
    Shared memory bank conflicts explainedWhat is a shared memory bank conflict, and why does adding one column fix it?
    run logged
  6. 16
    CUDA tiled matrix multiplication, measuredHow does tiling speed up matrix multiplication in CUDA?
    run logged
  7. 17
    Registers, local memory and spillsHow do I tell whether my CUDA kernel is spilling registers?
    run logged
  8. 18
    CUDA constant memory, measured against globalIs CUDA constant memory faster than global memory?
    run logged
  9. 19
    Unified memory and what it costsIs cudaMallocManaged slower than cudaMalloc?
    run logged
  10. 20
    CUDA image convolution and edge detectionHow do I write a fast 2D convolution in CUDA?
    run logged