MODULE 04 / Days 31-40
Parallel patterns
- 31CUDA prefix sum, the Hillis-Steele scanHow do I write a prefix sum in CUDA, and why does it need two shared arrays?run logged
- 32The CUDA scan algorithm: the tree and the three kernelsHow does a work-efficient CUDA scan work, and why three kernels?run logged
- 33CUDA stream compaction with a prefix sumHow do I remove elements from an array on a GPU?run logged
- 34CUDA stencils and the heat equationHow do I write a 2D stencil in CUDA, and why does my heat simulation blow up?run logged
- 35CUDA radix sort, built from the scan you already wroteHow does a radix sort work on a GPU, and why is it the sort GPUs use?run logged
- 36CUDA merge path: merging sorted arrays in parallelHow do you merge two sorted arrays in parallel on a GPU?run logged
- 37Sparse matrices in CUDA: COO, CSR and SpMVWhat is CSR, and why is my CUDA SpMV slow on some matrices?run logged
- 38Graph traversal on a GPUHow do you run a breadth-first search on a GPU, and why is one thread per vertex slow?run logged
- 39When to stop hand-writing: Thrust, CUB and libcu++Should I use Thrust and CUB or write the CUDA kernel myself?run logged
- 40PageRank on a real graph in CUDAHow do I implement PageRank in CUDA?run logged