← Glossary
CUDA glossaryPerformance
CC 7.5

What is memory bandwidth on a GPU?

Bytes per second between the SMs and DRAM, quoted as a theoretical peak on the box and measured as an effective figure by a real kernel.

The two numbers have two formulas, and every bandwidth conversation goes better once they are separate. The theoretical figure is hardware arithmetic: memory clock times bus width times transfers per clock, computed by the vendor and printed on the datasheet. For the Tesla T4, NVIDIA's Ada whitepaper appendix and the product page both say 320 GB/s (https://images.nvidia.com/aem-dam/Solutions/geforce/ada/nvidia-ada-gpu-architecture.pdf and https://www.nvidia.com/en-us/data-center/tesla-t4/ , both checked 2026-09-01). The effective figure is your arithmetic: bytes your kernel moved, reads plus writes, divided by measured time. No kernel reaches the spec, because refresh, command overhead and imperfect access streams all bill against the same pins.

The folklore says "expect 60 to 80 percent of the spec sheet", and the point of measuring is to replace folklore with a number for your card: this T4 delivers 76 percent to a well-behaved kernel. That measured ceiling, not the spec, is the honest denominator for every kernel you grade afterwards, which is why day 49 builds its roofline from the measured copy rather than the datasheet.

Effective bandwidth is also a property of the access pattern, not just the card. The same global memory served day 11's coalesced kernel at 243 GB/s and its stride-32 variant at 9.6 GB/s, a 25x collapse with identical hardware and identical bytes requested, because coalescing decides how many of the fetched bytes were wanted. When someone says "this kernel gets 12 GB/s", the first question is whether the card or the pattern is the limit, and only a measured ceiling can answer it.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75, captured 2026-09-01. Day 49 measured both ends of the card's roofline with its own kernels:

quantity value formula
theoretical bandwidth 320 GB/s vendor spec, FACT-SHEET section 3
measured copy, coalesced, in and out 244.7 GB/s (0.549 ms) 134,217,728 bytes / time
ratio 76 percent measured / spec

The same run measured the compute ceiling at 7179.2 GFLOP/s, putting the ridge point at 29.34 FLOP/byte: below that intensity nothing on this card can be compute bound, which is how one bandwidth measurement classifies every kernel you own. Nine of the ten kernels day 49 placed on this roofline sit left of the ridge, where this number, not the FLOP rate, is the wall.

Diagram

roofline-plotter, preset ten-kernels: the sloped ceiling drawn from the measured 244.7 GB/s, the flat ceiling from the measured GFLOP/s, and day 49's ten kernels plotted where their intensity and throughput landed.

Alt text: "A roofline built from the card's own measured bandwidth and compute ceilings, with nine of ten real kernels under the sloped memory limit."

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.