← Glossary
CUDA glossaryLibraries
CC any supported CUDA GPU

What is cuBLAS in CUDA?

NVIDIA's dense linear algebra library, providing architecture-tuned matrix operations whose data type, compute type and math mode define the numerical contract.

cuBLAS supplies GPU versions of BLAS operations, especially GEMM. It is the baseline for a hand-written matrix kernel because NVIDIA can choose a different implementation for each architecture, shape and number format. Calling it “the library result” is incomplete: input type, output type, accumulator type, transpose flags and math mode all affect both speed and error.

The traditional API inherits BLAS's column-major convention. Row-major matrices can still be multiplied by swapping operands and transpose flags, but that transformation must be written down or a correct call can look backwards. cublasGemmEx makes data and compute types explicit, which is essential when mixed precision or tensor-core paths are possible.

cuBLAS does not fuse arbitrary code around GEMM. A separate bias kernel writes and rereads the output. cuBLASLt exists for descriptor-based algorithm selection and epilogues such as bias. Day 81 holds the matrix arithmetic fixed and measures that packaging difference.

Measured

On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 81 measured a 2048-square FP32 cublasGemmEx followed by addBias at 6.1782 ms and 2,780.7 GFLOP/s. The bias kernel alone took 0.1357 ms at 247.4 GB/s. The matching cuBLASLt call with a separate bias took 6.1812 ms, only 0.05 percent different, and every path matched the CPU reference. These figures name FP32 input and CUBLAS_COMPUTE_32F; they are not TF32 or FP16 results.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.