What is cuSPARSE?
NVIDIA's sparse linear algebra library, providing descriptor-based operations such as CSR sparse matrix-vector multiply with algorithm and temporary-storage choices.
cuSPARSE implements operations whose work depends on a sparse structure rather than a dense rectangle. For CSR SpMV, you describe the matrix and dense vectors, query scratch-buffer size, choose an algorithm, then execute. The library has enough context to use a strategy beyond a fixed one-thread-per-row kernel.
That matters when row lengths are uneven. A thread-per-row kernel leaves lanes idle on long, skewed rows; a warp-per-row kernel wastes work on short even rows. cuSPARSE can outperform both, but its descriptor setup and scratch allocation should be amortized outside the timed loop. Algorithm choice also affects floating-point determinism: ALG2 is documented for bitwise repeatability, which may matter more than a small timing difference.
cuSPARSE and cuSPARSELt are related but not interchangeable. The latter targets structured tensor-core sparsity such as 2:4. Day 98's T4 cannot run that sm_80 path, so its structured-sparsity result is accuracy-only.
Measured
On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 82 measured skewed CSR SpMV. cuSPARSE ALG1 took 0.366 ms at 45.8 GFLOP/s; the best hand kernel, warp per row, took 1.196 ms at 14.0 GFLOP/s, so the library was 3.27x faster. On the even matrix, ALG2 took 0.379 ms at 44.2 GFLOP/s, 1.61x faster than the best hand row. Scratch requirements were 32,772 bytes for ALG1 and 164,484 bytes for ALG2. All four paths matched the independent reference.
Diagram: one even CSR matrix and one skewed matrix feed thread-per-row, warp-per-row and cuSPARSE paths. The takeaway is that row-length distribution, not only nonzero count, chooses the fastest mapping.
Related terms
Where you meet this
- Day 82, CUDA math libraries, which produced the same-process comparison.
- Day 37, sparse matrices, where both hand-written mappings are built.
- How to set up CUDA, because cuSPARSE is linked separately from the runtime.
Sources
- cuSPARSE documentation, generic SpMV and algorithms: https://docs.nvidia.com/cuda/cusparse/index.html (checked 2026-09-01)
- cuSPARSELt documentation, structured sparse operations: https://docs.nvidia.com/cuda/cusparselt/index.html (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.