← Glossary
CUDA glossaryPerformance
CC 7.5

What is latency hiding in CUDA?

Keeping enough warps in flight that the SM always has one ready to issue while the others wait on memory.

A global load takes hundreds of cycles, and a GPU does not fight that with a big cache and out-of-order tricks the way a CPU core does. It keeps many warps resident instead. When one warp issues a load and stalls, the warp scheduler picks another that is ready, at no switching cost, because every resident warp's registers already sit on the SM. The latency did not shrink; it got covered by other work. That is the whole mechanism, and occupancy is how many candidates the scheduler has.

What most occupancy discussions skip is that warps are only one source of in-flight work. A single thread that issues four independent loads before using any of them has four requests outstanding, which covers latency exactly as well as four threads issuing one each. This is instruction-level parallelism, and it is why "raise occupancy" is not the only fix for a latency-bound kernel, and why two kernels at very different occupancies can reach the same speed. The bookkeeping identity is Little's law: concurrency equals latency times throughput, so a fixed bandwidth target at a fixed latency needs a fixed number of requests in flight, and the hardware does not care whether warps or ILP supply them.

The practical reading: when a kernel is slow and the profiler shows long-scoreboard stalls, the question is not "is occupancy high" but "how many independent requests does each thread put in flight, times how many threads are resident". Raise either factor until the memory system saturates; past that point, both are free to fall.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75, captured 2026-09-01. Day 45 held the in-flight load count at 64 while moving occupancy, then held occupancy while cutting the in-flight count ("in flight" is resident warps times the loads one thread issues before its first multiply):

label occupancy loads in flight ms GB/s
tlp1 100% 64 0.785 256.5
ilp4-b1 25% 64 0.798 252.2
tlp1-b1 25% 16 1.139 176.7

Four times less occupancy at the same concurrency cost a ratio of 1.02. The same occupancy at a quarter of the concurrency cost 0.69 of the best bandwidth. Concurrency is the variable; occupancy is one way to buy it. The ILP rows paid for their extra in-flight loads in registers (32 per thread for four items against 10 for one), which is the real trade: ILP spends register budget where occupancy spends warp slots.

Diagram

occupancy-stepper, preset two-ways-to-64: two SMs side by side, one with 32 resident warps issuing one load each, one with 8 resident warps issuing four, both showing 64 requests outstanding.

Alt text: "Two configurations keep 64 loads in flight, one with many warps and one with few warps doing more independent work each. Measured on a T4, they run within 1.02x of each other."

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.