What is the CUDA runtime API?
The cuda* calls in cudart that almost all CUDA code uses, layered on top of the driver API.
You get it from cuda_runtime.h and libcudart, both of which ship with the toolkit, so its version is the toolkit's version. What it buys you is everything you never wrote: NVIDIA describes the runtime as giving "implicit primary context initialization and management, and implicit module management", against a driver API where "the execution configuration and kernel parameters must be specified with explicit function calls". That is the difference between kernel<<<blocks, threads>>>(p) and a page of setup. Every call hands back a cudaError_t, which is why error checking is a habit rather than an option, and why an unchecked program can run to completion having done nothing.
The one thing the toolkit cannot give you is the driver. Your binary carries a runtime from the toolkit that built it, the driver arrives separately from a package or a cloud image, and the driver has to be at least as new as the runtime asks for.
| What you build with | Driver it needs |
|---|---|
| any CUDA 13.x, minor version compatibility (Table 2) | 580 or newer |
| CUDA 13.3 Update 1 | 610.43.02 or newer |
Go under the floor and you get cudaErrorInsufficientDriver, error 35: CUDA driver version is insufficient for CUDA runtime version. It is not a broken install. It is a new toolkit on an old driver, and the fix is the driver. The two numbers do not have to match, which is the part most pages state wrongly: the release notes carry two different floors and they are easy to confuse. Table 2 gives 580 as the minor version compatibility floor for the 13.x line. Table 3 gives the toolkit driver version per release, and that is the one that binds: CUDA 13.3 GA and 13.3 Update 1 both need 610.43.02.
Two details catch people who thought they knew this API. Since CUDA 12.0, cudaRuntimeGetVersion "As of CUDA 12.0, this function no longer initializes CUDA" and its purpose "is solely to return a compile-time constant stating the CUDA Toolkit version", so it tells you what compiled the binary and nothing about the machine it is on. And the documented behaviour that "all the kernels are automatically loaded during initialization and stay loaded for as long as the program runs" is no longer the default: lazy module loading has been on since 12.2, which is why every kernel you intend to time needs its own warm-up.
Measured
On a Tesla T4 (driver 595.84, CUDA 12.6, V12.6.85, built with nvcc -O3 -arch=sm_75), day 3 printed runtime version 12.6 and max CUDA the driver takes 13.2 side by side, captured 2026-08-30 on the project's verification node. The runtime is older than the ceiling, so the program runs. Reverse the two and you get error 35. That gap is also why the course's own code is verified on 12.6: this node's driver will not accept the 13.3 Update 1 runtime. Full run behind How to set up CUDA.
Diagram
Related terms
CUDA driver API · nvcc · CUDA error checking · cudaMemcpy · lazy module loading · compute capability
Where you meet this
Day 3, how to set up CUDA, reads both version numbers off a real card; which CUDA version works with your driver turns the table above into a decision. Day 5 uses the API for real, allocating, copying and launching with a check on every call. When the driver is behind, the page is CUDA driver version is insufficient for CUDA runtime version.
Sources
- Implicit context and module management, and what the driver API makes explicit: https://docs.nvidia.com/cuda/cuda-driver-api/driver-vs-runtime-api.html (checked 2026-08-30)
cudaRuntimeGetVersionreturning a compile-time constant since 12.0: https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART____VERSION.html (checked 2026-08-30)- Driver floors, Tables 2 and 3: https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html#cuda-driver (checked 2026-08-29)
- Lazy loading as the default: https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/lazy-loading.html (checked 2026-08-29)
Byline
Author and reviewer are unassigned. This entry publishes when two different named people have signed it; the written and last-checked dates are set then.