How many warps and blocks can one SM hold?
The hard caps on how many warps and blocks one SM can hold at once, which set the ceiling occupancy can reach.
An SM holds whole blocks, and everything a resident block needs, registers for every thread, its shared memory, a warp slot per 32 threads, is allocated for as long as the block lives. Two of the limits are architectural constants: a cap on resident warps and a cap on resident blocks, whichever binds first. The rest are resource arithmetic against the register file and shared memory. Occupancy is achieved warps over the warp cap, so the cap is the denominator of the whole discussion.
The trap is assuming the constants are universal. Table 30 of the compute capabilities appendix says otherwise: 32 resident warps per SM at compute capability 7.5, 64 at 8.0 and 9.0, and 48 at 8.9 and 12.0, with resident blocks ranging from 16 to 32 (https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/compute-capabilities.html , checked 2026-08-29). The same 256-thread block is a quarter of a T4's warp budget and an eighth of an H100's, so a hardcoded occupancy calculator is wrong on somebody's card by construction. This is why day 45 asks the driver (cudaOccupancyMaxActiveBlocksPerMultiprocessor) rather than the tables: the driver also knows the allocation granularities the published tables omit.
One more consequence worth naming: the block cap means small blocks can strand warp slots. On the T4, 32-thread blocks hit the 16-block cap at 16 warps, half the 32-warp budget, so 50 percent occupancy is the most such a launch can ever reach. Day 2 predicted exactly that before running, and the measurement held. The caps interact, and the binding one is a property of your launch shape, not just your card.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75, captured 2026-09-01. Day 45 read the limits off the device, 32 resident warps and 16 resident blocks per SM, 65536 registers and 65536 bytes of shared memory per SM, then walked residency down with a never-read dynamic shared-memory reservation on a 256-thread (8-warp) block:
| reservation (B) | blocks/SM (driver) | warps/SM | occupancy |
|---|---|---|---|
| 0 | 4 | 32 | 100% |
| 21760 | 3 | 24 | 75% |
| 32768 | 2 | 16 | 50% |
| 49152 | 1 | 8 | 25% |
At zero reservation the warp cap binds (4 blocks of 8 warps is exactly 32); each larger reservation makes shared memory the binding limit instead. Every row's blocks-per-SM figure is the driver's answer, and the measured effect of walking this ladder on the kernel's speed is the occupancy entry's table.
Diagram
occupancy-stepper, preset residency-dials: one SM's 32 warp slots filling and emptying as the shared-memory reservation steps through the four rows above.
Alt text: "A T4 SM holds at most 32 warps. Four 8-warp blocks fill it; growing each block's shared memory request evicts one resident block at a time."
Related terms
Where you meet this
- Day 45, occupancy, the lesson that owns this term and walked the ladder above.
- Day 2, how a GPU differs from a CPU, where the caps first appear and a prediction about which binds is tested.
- Day 17, registers and spills, the register file's version of the same arithmetic.
Sources
- CUDA Programming Guide, compute capabilities appendix, Table 30, for the per-architecture resident warp and block limits: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/compute-capabilities.html (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.