← Glossary
CUDA glossaryHardware
CC any

What is a streaming multiprocessor?

The unit a GPU is actually made of: it owns a slice of registers and shared memory, holds many blocks at once, and issues warps from whichever one is ready.

A GPU is a bag of SMs. Each one has its own register file, its own shared memory, a few warp schedulers and a set of arithmetic pipelines, and none of that is shared with the SM next door. A thread block lands on one SM and stays there for its whole life, which is why __syncthreads() is block-wide and why a block whose registers or shared memory do not fit on one SM is not moved somewhere roomier. It fails to launch, with too many resources requested for launch.

What an SM hands out is warp slots, not threads. A block of T threads costs ceil(T / 32) slots and the division rounds up, so a 48-thread block takes two slots and runs the second one with 16 of its 32 lanes switched off for as long as the block lives. Four caps then decide how many blocks sit on the SM at once: resident warps, resident blocks, the register file, and shared memory. Whichever binds first is your answer, and it moves with the architecture rather than with the card's price: 32 resident warps per SM at compute capability 7.5, 48 on an RTX 50, 64 on an H100 (Table 30 of the compute capabilities appendix).

The mistake worth naming is reading an SM as a CPU core and the CUDA cores inside it as its hyperthreads. An SM does not hold one thread of control that it saves and restores. It holds the registers and program counters of every resident warp on chip at once, so switching between warps costs nothing and the scheduler issues from whichever warp is ready this cycle. That is the point of occupancy: it is the supply of warps the scheduler gets to choose from, not a score to maximise.

Measured

On a Tesla T4 (driver 595.84, CUDA 12.6, built with nvcc -O3 -arch=sm_75), day 2 asked the driver instead of a table and got 40 SMs, each holding 32 resident warps, 16 resident blocks, 65,536 32-bit registers and 64 KiB of shared memory, in front of a 4096 KiB L2 for the whole device. Shared memory per block is 48 KiB by default on this card, and day 3 measured the 64 KiB opt-in.

Sweeping block sizes through cudaOccupancyMaxActiveBlocksPerMultiprocessor shows which cap binds. At 16 and at 32 threads per block the block cap goes first: 16 blocks fit, they fill 16 of the 32 warp slots, and the SM sits at 50 percent occupancy no matter how many blocks you launch. At 48, 64, 128, 256, 512 and 1024 threads the warp slots fill and occupancy reaches 100 percent. It does not follow that every size above 48 does: day 10 swept the same card and found 96, 160 and 192 threads at 94 percent and 384 and 768 at 75, because a block size that does not divide the 32-warp budget cleanly leaves slots stranded. Captured 2026-08-30 on the project's verification node.

Alt text for the static frame: one SM with 32 warp slots, filled by 16 blocks of 32 threads to half its height, and by 16 blocks of 48 threads to the top.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.