What is cuDNN's graph API?
NVIDIA's deep-learning primitive library, whose graph API turns described operations into a selected engine, execution plan, workspace and bound tensor addresses.
The graph API separates what to compute from how to compute it. Tensor descriptors carry dimensions, strides, data types and unique IDs; operation descriptors connect tensors to work such as a convolution; an operation graph groups those operations. No device address is needed while that graph is being built.
A heuristics query ranks engine configurations for the graph. The application tries configurations until one finalizes as an execution plan, then allocates the plan's requested workspace. At execution, a variant pack pairs tensor IDs with device pointers and passes them to cudnnBackendExecute. This model can represent fused graphs that the deprecated one-function-per-operation API could not express.
Header and runtime versions must be checked independently. CUDNN_VERSION describes the headers used at compile time, while cudnnGetVersion() reports the shared object loaded at run time. PyTorch wheels bundle their own CUDA and cuDNN builds, so a separate PyTorch process may use the same convolution semantics without loading the standalone program's exact library version.
Measured
On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 83 ran an FP32 graph for x[2,64,56,56] * w[64,64,3,3]. The standalone header and runtime both reported cuDNN 91301. Heuristics mode A returned 8 configurations; rank 0 finalized as engine global index 12 with 409856 bytes of workspace.
The largest difference from the float64 Kahan reference was 1.268e-05, within the program's derived tolerance. A separate torch 2.13.0+cu126 process, bundling cuDNN 9.10.2, matched all 401408 standalone outputs with maximum difference 0.000e+00. Day 83 deliberately records no execution time, so this page makes no cuDNN performance claim.
Related terms
Where you meet this
- Day 83, cuDNN with the graph API, which builds, executes and checks the convolution graph above.
- Day 81, cuBLAS and cuBLASLt, where another library exposes heuristics and workspace choices for matrix multiplication.
Sources
- cuDNN Backend API overview: https://docs.nvidia.com/deeplearning/cudnn/backend/latest/api/overview.html (checked 2026-09-01)
- Verifying the loaded cuDNN installation: https://stackoverflow.com/questions/31326015/how-to-verify-cudnn-installation (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified evidence was captured on 2026-09-02; publication still requires named author and reviewer sign-off.