← Glossary
CUDA glossaryHardware
CC 7.5

What are GPU architecture generations?

Turing, Ampere, Ada, Hopper and Blackwell, each mapping to a range of compute capabilities and a set of features.

The generation is the silicon family and marketing name; the compute capability is its version number. You need both, because documentation, tuning guides and release notes talk in names while every compiler flag and feature table talks in numbers.

Generation Compute capability Card this course names What it is here for
Turing 7.5 Tesla T4 The floor. Colab free, Kaggle T4 x2, and the Compiler Explorer runner
Ampere 8.0 A100 cp.async and hardware async barriers
Ada Lovelace 8.9 L4, RTX 4090 Cheap FP8, the common home card
Hopper 9.0 H100 TMA, clusters, wgmma
Blackwell, datacenter 10.0 B200 tcgen05, concept only
Blackwell, consumer 12.0 RTX 5090 Clusters and single-CTA TMA, no wgmma

Read the last two rows twice. Blackwell is two machines. A GeForce RTX 50 is compute capability 12.0 and a B200 is 10.0, so the consumer card carries the higher number and the smaller feature set. Neither one is a superset of the other, and no generation name on its own tells you whether an instruction exists. Check the row in compute capability before you write against a feature.

The other half of this term is the drop list. CUDA 13.0 removed offline compilation and library support for Maxwell, Pascal and Volta, so anything below Turing no longer builds with a current toolkit. nvcc's default target moved with it: "sm_75 is used as the default value; PTX is generated for compute_75 then assembled and optimized for sm_75". Cards below the line are not bricked. The 12.x toolkits still build for them, and R580 is the last driver branch that carries those binaries. What you cannot do is put a modern toolkit on an old card, which is exactly what happens on Kaggle's free default: a Tesla P100 is 6.0, and nvcc answers nvcc fatal : Unsupported gpu architecture 'sm_60'. Pick T4 x2 instead, per free GPU tiers.

Generations also move the constants your kernels are budgeted against, and those move in both directions. A Turing SM holds 32 resident warps; Hopper holds 64 and consumer Blackwell holds 48. Shared memory per thread block goes 64 KB on a T4, 163 KB on an A100, 227 KB on an H100, and back down to 99 KB on an RTX 5090. Any tile size or occupancy argument you carry across a generation boundary has to be recomputed, not reused.

Measured

On a Tesla T4 (driver 595.84, CUDA 12.6, V12.6.85, built with nvcc -O3 -arch=sm_75), day 3 read Turing's own numbers off the card: compute capability 7.5, shared memory per block 48 KiB with 64 KiB available as an opt-in, L2 4096 KiB, 65536 32-bit registers per block, 40 SMs, and 1024 max resident threads per SM. Captured 2026-08-30 on the project's verification node. The 48 against 64 is the Turing detail worth keeping: the default cap per block is not the SM's capacity, and reaching the rest needs an explicit cudaFuncSetAttribute. Full run behind How to set up CUDA.

Related terms

compute capability · tensor core · gencode · thread block cluster · free GPU tiers · streaming multiprocessor

Where you meet this

Day 3, how to set up CUDA, is where the generation decides whether your toolkit will talk to your card at all, and day 0 uses it to choose the hardware. The name comes back whenever a feature does: Ampere on day 74 for cp.async, Hopper on day 76 for TMA, and consumer Blackwell on day 78. The error a wrong pairing gives is nvcc fatal: Unsupported gpu architecture.

Sources

Byline

Author and reviewer are unassigned. This entry publishes when two different named people have signed it; the written and last-checked dates are set then.