← Glossary
CUDA glossaryLibraries
CC any supported CUDA GPU

What are the CUDA math libraries?

Toolkit libraries that provide tuned random generation, transforms, sparse algebra and dense solvers, replacing custom kernels while exposing plans, state and scratch-buffer costs.

CUDA ships specialized libraries for algorithms that are expensive to implement and tune. cuRAND generates random streams, cuFFT plans and executes transforms, cuSPARSE handles sparse algebra, and cuSOLVER factors and solves dense systems. cuBLAS covers dense BLAS separately.

One library call can hide several costs. cuRAND's device API initializes state per thread; its host API fills global memory for another kernel to reread. cuFFT builds a plan and work area. cuSPARSE queries scratch bytes before execution. cuSOLVER reports factorization status through device memory. Setup belongs in a separate table from repeated work, or the benchmark answers the wrong lifecycle question.

Correctness contracts also differ. cuFFT is unnormalized, so a forward followed by inverse transform returns N times the input. Random estimates need statistical bounds rather than elementwise equality. A library should be checked against an independent analytic or CPU reference, never against a second call to itself.

Measured

On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 82 measured a batched set of 1,024 FFTs at 0.074 ms, versus 5.540 ms for 1,024 batch-one calls, a 74.77x ratio. The cuRAND device path spent 0.042 ms initializing state and 0.473 ms sampling; its 0.515 ms total beat the host API's 1.292 ms. cuSOLVER getrf took 14.180 ms and getrs 3.851 ms. The FFT round trip measured scale 1023.9999, confirming the documented missing 1/N normalization. All corrected-run checks passed.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.