MODULE 04 / Days 31-40

Parallel patterns

All lessons
  1. 31
    CUDA prefix sum, the Hillis-Steele scanHow do I write a prefix sum in CUDA, and why does it need two shared arrays?
    run logged
  2. 32
    The CUDA scan algorithm: the tree and the three kernelsHow does a work-efficient CUDA scan work, and why three kernels?
    run logged
  3. 33
    CUDA stream compaction with a prefix sumHow do I remove elements from an array on a GPU?
    run logged
  4. 34
    CUDA stencils and the heat equationHow do I write a 2D stencil in CUDA, and why does my heat simulation blow up?
    run logged
  5. 35
    CUDA radix sort, built from the scan you already wroteHow does a radix sort work on a GPU, and why is it the sort GPUs use?
    run logged
  6. 36
    CUDA merge path: merging sorted arrays in parallelHow do you merge two sorted arrays in parallel on a GPU?
    run logged
  7. 37
    Sparse matrices in CUDA: COO, CSR and SpMVWhat is CSR, and why is my CUDA SpMV slow on some matrices?
    run logged
  8. 38
    Graph traversal on a GPUHow do you run a breadth-first search on a GPU, and why is one thread per vertex slow?
    run logged
  9. 39
    When to stop hand-writing: Thrust, CUB and libcu++Should I use Thrust and CUB or write the CUDA kernel myself?
    run logged
  10. 40
    PageRank on a real graph in CUDAHow do I implement PageRank in CUDA?
    run logged