What is the register file on a GPU?
The fixed pool of registers on each SM, partitioned among every thread resident there, which is what caps occupancy in most real kernels.
Caches share. The register file does not. When a block lands on an SM, the hardware carves out that block's registers and holds them until the block finishes, so nothing is evicted, nothing is refilled, and there is no hit rate to improve. The consequence is a hard division rather than a soft one: whatever the file does not have room for cannot be resident, and the SM runs with empty warp slots for the whole launch.
Do the division in blocks, because a block is what the hardware allocates. A T4 SM holds 65,536 32-bit registers and up to 1,024 threads, which puts the break-even at 64 registers a thread. Launch 256 threads a block and a kernel using 72 registers needs 18,432 registers per block, so three blocks fit and 768 threads are resident. Cut the same kernel to 64 and a block costs 16,384, four fit, and all 1,024 slots are used. Nothing gradual happens in between. Going from 64 registers to 65 costs a whole block, which is a quarter of the SM.
The number people misread here is which side of the fraction is per SM. The 65,536 is a per-SM figure from Table 31 of the compute-capability appendix, not a per-block or per-kernel one, and it changes with compute capability rather than with your card's price. Registers are also not the only per-SM budget: shared memory is partitioned the same way, and occupancy is the smallest answer the three limits give, not the register answer alone. Which limit is binding is a fact about your kernel, and the only honest way to find it is to read all three.
Measured
The card reports its own pool before anything runs: 65,536 32-bit registers per SM and 1,024 resident threads per SM. Day 17 then ran two versions of one kernel against it at 256 threads a block, on a Tesla T4 with driver 595.84 and CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75.
| Kernel | Registers per thread | Registers per block, arithmetic | Blocks per SM | Occupancy |
|---|---|---|---|---|
decayWide |
72 | 18,432 | 3 | 75.0 % |
decayWideBounded |
64 | 16,384 | 4 | 100.0 % |
The register counts, the block counts and the occupancy came from the device through cudaFuncGetAttributes and cudaOccupancyMaxActiveBlocksPerMultiprocessor. The middle column is the multiplication you can do yourself, and it lands where the runtime does, which is the point: this is a budget you can settle before you compile, not a profiler finding. What the arithmetic does not tell you is whether the 100 percent version is faster. On this pair it was not, and register pressure has that measurement.
Diagram
occupancy-stepper, register axis. The SM's 65,536 registers drawn as a bar that fills in 256-thread block units, next to the 1,024 thread slots those blocks occupy. Stepping the register count from 64 to 72 drops the fourth block out of both bars at once.
Alt text: "A T4 SM's register file as a bar filled in whole blocks. At 64 registers a thread, four 256-thread blocks fit and every thread slot is used. At 72 the fourth block does not fit and a quarter of the SM stands idle."
Related terms
- registers
- register pressure
- register spilling
- occupancy
- streaming multiprocessor
- resident warps per SM
- shared memory
Where you meet this
- Day 17, registers, local memory and spills, the lesson that owns this term and does the division against four kernels.
- Day 10, choosing threads per block, where the block size decides how much of the file each block claims.
- Day 2, how a GPU differs from a CPU, for why a chip carries a register file this large in the first place.
- How to set up CUDA, whose device query prints your own card's
regsPerMultiprocessor. too many resources requested for launch, the error a launch returns when one block's demand does not fit the pool.
Sources
- CUDA Programming Guide, Table 30 "Device and Streaming Multiprocessor (SM) Information per Compute Capability", for registers per SM and resident threads per SM: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/compute-capabilities.html (checked 2026-08-30)
- CUDA C++ Best Practices Guide, 10.2.7.1 "Register Pressure", on registers as a shared, limited resource: https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html (checked 2026-08-30)
- Turing Tuning Guide, for the per-SM resources on compute capability 7.5: https://docs.nvidia.com/cuda/turing-tuning-guide/index.html (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Two different named people have to sign this entry before it leaves draft.