What is compute capability in CUDA?
The version number that says which features a GPU has, written 7.5 in docs and sm_75 on the nvcc command line.
It is a property of the silicon, not of anything you install. A Tesla T4 is 7.5 forever; an A100 is 8.0, an RTX 5090 is 12.0. The runtime reads it back as major and minor, the docs write it with a dot, and nvcc wants the dot gone, which is where -arch=sm_75 comes from. Two other numbers get mistaken for it: the toolkit nvcc --version prints, and the highest CUDA runtime your driver accepts, which is what the nvidia-smi header shows. Day 3 prints all three together so you can see which one moved.
The number is a feature gate, and the row matters more than the marketing name of the generation.
| Feature | 7.5 | 8.0 | 8.9 | 9.0 | 10.0 | 12.0 |
|---|---|---|---|---|---|---|
cp.async |
no | yes | yes | yes | yes | yes |
| FP8 | no | no | yes | yes | yes | yes |
| Thread block clusters | no | no | no | yes | yes | yes |
| Distributed shared memory | no | no | no | yes | yes | yes |
| TMA, single CTA | no | no | no | yes | yes | yes |
| TMA multicast, in hardware | no | no | no | yes | yes | no |
wgmma |
no | no | no | sm_90a only |
no | no |
| tcgen05 | no | no | no | no | yes | no |
The bold cells are the ones most pages get wrong. A consumer Blackwell card at 12.0 does have clusters, distributed shared memory and single-CTA TMA. It has no wgmma, which the PTX target notes pin to sm_90a and which therefore exists on no Blackwell at all, and no tcgen05. A bigger number is not a superset either: 12.0 is a GeForce card and 10.0 is a B200.
Get the number wrong in either direction and you get one of two failures. Ask for something under your toolkit's floor and nvcc stops before it compiles anything: nvcc fatal : Unsupported gpu architecture 'sm_60', with three spaces before the colon, which matters when you paste it into a search box. Build above your card instead, -arch=sm_80 on a T4, and it compiles, links and starts, then the first launch returns error 209, no kernel image is available for execution on the device. -arch=sm_75 survives newer cards because it also emits compute_75 PTX for the driver to JIT. -arch=native does not: nvcc generates code only for the GPUs it can see, and "no PTX program will be generated for this option".
Measured
On a Tesla T4 (driver 595.84, CUDA 12.6, V12.6.85, built with nvcc -O3 -arch=sm_75), day 3's devicequery read compute capability 7.5 off the device and 750 out of the binary, captured 2026-08-30 on the project's verification node. Same number, two notations: major.minor from cudaGetDeviceProperties, and the three-digit __CUDA_ARCH__ nvcc defines during the device pass. Full run behind How to set up CUDA.
Code
The card's number is easy to read. The binary's number only exists inside device code, so a kernel has to write it out:
__global__ void reportCompiledArch(int* out, size_t n) {
const size_t i = blockIdx.x * static_cast<size_t>(blockDim.x) + threadIdx.x;
if (i < n) {
#if defined(__CUDA_ARCH__)
out[i] = __CUDA_ARCH__;
#else
out[i] = 0;
#endif
}
}
The #else branch is not dead code. The same file is compiled once for the host, where __CUDA_ARCH__ is undefined.
Related terms
gencode · GPU architecture generations · nvcc · architecture-specific target · JIT compilation · free GPU tiers
Where you meet this
Day 3, how to set up CUDA, reads it off your card and picks the flag. Day 0 uses it to choose hardware, day 69 revisits it when one binary must serve several cards, and day 78 lands on the corrected row above. Both failures have pages: nvcc fatal: Unsupported gpu architecture and no kernel image is available for execution on the device.
Sources
- Feature support per compute capability, Table 29: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/compute-capabilities.html (checked 2026-08-29)
-archshorthand, thesm_75default andnative: https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html#gpu-architecture-arch (checked 2026-08-30)- Which capability your card has: https://developer.nvidia.com/cuda-gpus (checked 2026-08-29)
wgmmarequiressm_90a: https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#asynchronous-warpgroup-level-matrix-instructions-wgmma-mma (checked 2026-08-29)- No multicast on GeForce, so the cluster shape is fixed to 1x1x1: https://github.com/NVIDIA/cutlass/blob/main/media/docs/cpp/blackwell_functionality.md#cluster-size (checked 2026-08-29)
Byline
Author and reviewer are unassigned. This entry publishes when two different named people have signed it; the written and last-checked dates are set then.