What is Little's law in GPU programming?
Concurrency equals latency times throughput, which is why you need many memory requests in flight to saturate a GPU's bandwidth.
The law comes from queueing theory and it is bookkeeping, not physics: if each request takes L seconds and you want T requests completing per second, then L times T requests must be in the system at any moment. Nothing about GPUs in that sentence, which is its power. A global load's latency is hundreds of cycles and the bandwidth target is hundreds of gigabytes per second, so the concurrency the product demands is large, and the entire SM design, thousands of resident threads, zero-cost warp switching, follows from having to hold that concurrency somewhere.
What the law predicts, and what makes it more than a slogan, is that concurrency is the only variable. It does not care whether the in-flight requests come from many warps each holding one, or few warps each holding several independent ones. Occupancy and instruction-level parallelism are two suppliers of the same commodity, so trading one for the other at constant product should not change the speed, and starving the product should. Both predictions are checkable on real hardware, which is what day 45 does. NVIDIA's own tuning advice frames occupancy in exactly these terms, as a means of keeping the machine's latency covered rather than a target (https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html#occupancy , checked 2026-08-30).
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75, captured 2026-09-01. Day 45 counted "loads in flight" per SM as resident warps times independent loads issued per thread before the first use, then moved the two factors against each other:
| label | occupancy | loads in flight | ms |
|---|---|---|---|
| tlp1 | 100% | 64 | 0.785 |
| ilp4-b1 | 25% | 64 | 0.798 |
| tlp1-b1 | 25% | 16 | 1.139 |
Held product, swapped supplier: a ratio of 1.02, the law's first prediction. Cut the product to a quarter: 0.69 of the best bandwidth, the second. The constant of proportionality is the load latency, which this run does not measure separately; what it verifies is the law's structure, that the product is what the hardware responds to. Eight warps carrying eight loads each behaved like thirty-two carrying two.
Related terms
Where you meet this
- Day 45, occupancy, the lesson that owns this term and ran the exchange above.
- Day 49, the measured roofline, where the bandwidth the law must feed is measured.
- Day 2, how a GPU differs from a CPU, the design consequence of the law, met before the law itself.
Sources
- CUDA C++ Best Practices Guide, "Occupancy", for latency coverage as the stated purpose of resident warps: https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html#occupancy (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.