← Glossary
CUDA glossaryPerformance
CC 7.5

What is the roofline model?

A plot of achievable throughput against arithmetic intensity, with a sloped memory limit and a flat compute limit, that tells you which one you are hitting.

The x axis is arithmetic intensity, FLOP per byte. The y axis is throughput. Two ceilings bound everything: a sloped line where performance equals intensity times memory bandwidth, and a flat line at peak FLOP rate. They cross at the ridge point, and a kernel's position against it is the diagnosis: left of the ridge, memory-bound, and faster math buys nothing; right of it, compute-bound, and better memory access buys nothing. One chart replaces an argument.

A roofline is only as honest as its inputs, and both can lie. Vendor ceilings first: day 49 measures its own instead, a copy kernel for the slope and an FMA chain for the flat line, because the achievable ceiling on this T4 is 244.7 GB/s, not the 320 on the datasheet. Then the bytes: intensity computed from the bytes the algorithm asks for silently assumes every byte comes from DRAM. Day 30 placed a naive matmul at an impossible 1008.4 percent of the copy ceiling exactly this way, and documented the flaw rather than hiding it. The fix is the profiler's dram__bytes counters, what DRAM actually moved, which is the number the ratio was always supposed to have.

Placed with measured bytes, the same naive matmul lands at 3.8 percent of the ceiling with a DRAM-to-ask ratio of 0.0039: its re-reads never left cache, so it is neither DRAM-bound nor compute-bound but latency-and-cache-bound, a state the two-ceiling picture cannot name and the follow-up in Nsight Compute can. That is the model's real limit: it classifies kernels against DRAM and FLOP walls, and everything in between belongs to the profiler.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75, captured 2026-09-01. Day 49 measured both ends, then placed ten course kernels twice, first with asked-for bytes, then with dram__bytes from Nsight Compute:

copy, coalesced, in and out    0.549 ms    244.7 GB/s
fma, 8 chains                  1.496 ms   7179.2 GFLOP/s
ridge point                                29.34 FLOP/byte
kernel ask-based %ceil DRAM-based %ceil DRAM/ask
vector add 104.6 116.9 1.0922
vector add, stride 32 4.8 107.6 21.5566
matmul, naive 1008.4 3.8 0.0039
matmul, tiled 16 101.8 5.7 0.0579

The strided add's two rows tell the same story from the other side: it asks for 201.3 MB and DRAM moves 21.6 times that, so against the bytes it actually moved it runs at full ceiling; the waste is in what it moved, not how fast. Nine of the ten kernels sit left of the ridge, which is the course's whole memory-first argument drawn as one chart.

Diagram

roofline-plotter, preset ten-kernels: the two measured ceilings with all ten kernels plotted, and a toggle between ask-based and DRAM-based placement.

Alt text: "A roofline from the card's own measured ceilings. Toggling from asked-for bytes to measured DRAM bytes moves the naive matmul from an impossible 1008 percent of ceiling to 4 percent, inside the cache."

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.