What is the roofline model?
A plot of achievable throughput against arithmetic intensity, with a sloped memory limit and a flat compute limit, that tells you which one you are hitting.
The x axis is arithmetic intensity, FLOP per byte. The y axis is throughput. Two ceilings bound everything: a sloped line where performance equals intensity times memory bandwidth, and a flat line at peak FLOP rate. They cross at the ridge point, and a kernel's position against it is the diagnosis: left of the ridge, memory-bound, and faster math buys nothing; right of it, compute-bound, and better memory access buys nothing. One chart replaces an argument.
A roofline is only as honest as its inputs, and both can lie. Vendor ceilings first: day 49 measures its own instead, a copy kernel for the slope and an FMA chain for the flat line, because the achievable ceiling on this T4 is 244.7 GB/s, not the 320 on the datasheet. Then the bytes: intensity computed from the bytes the algorithm asks for silently assumes every byte comes from DRAM. Day 30 placed a naive matmul at an impossible 1008.4 percent of the copy ceiling exactly this way, and documented the flaw rather than hiding it. The fix is the profiler's dram__bytes counters, what DRAM actually moved, which is the number the ratio was always supposed to have.
Placed with measured bytes, the same naive matmul lands at 3.8 percent of the ceiling with a DRAM-to-ask ratio of 0.0039: its re-reads never left cache, so it is neither DRAM-bound nor compute-bound but latency-and-cache-bound, a state the two-ceiling picture cannot name and the follow-up in Nsight Compute can. That is the model's real limit: it classifies kernels against DRAM and FLOP walls, and everything in between belongs to the profiler.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75, captured 2026-09-01. Day 49 measured both ends, then placed ten course kernels twice, first with asked-for bytes, then with dram__bytes from Nsight Compute:
copy, coalesced, in and out 0.549 ms 244.7 GB/s
fma, 8 chains 1.496 ms 7179.2 GFLOP/s
ridge point 29.34 FLOP/byte
| kernel | ask-based %ceil | DRAM-based %ceil | DRAM/ask |
|---|---|---|---|
| vector add | 104.6 | 116.9 | 1.0922 |
| vector add, stride 32 | 4.8 | 107.6 | 21.5566 |
| matmul, naive | 1008.4 | 3.8 | 0.0039 |
| matmul, tiled 16 | 101.8 | 5.7 | 0.0579 |
The strided add's two rows tell the same story from the other side: it asks for 201.3 MB and DRAM moves 21.6 times that, so against the bytes it actually moved it runs at full ceiling; the waste is in what it moved, not how fast. Nine of the ten kernels sit left of the ridge, which is the course's whole memory-first argument drawn as one chart.
Diagram
roofline-plotter, preset ten-kernels: the two measured ceilings with all ten kernels plotted, and a toggle between ask-based and DRAM-based placement.
Alt text: "A roofline from the card's own measured ceilings. Toggling from asked-for bytes to measured DRAM bytes moves the naive matmul from an impossible 1008 percent of ceiling to 4 percent, inside the cache."
Related terms
Where you meet this
- Day 49, the measured roofline, the lesson that owns this term and built the chart.
- Day 30, the memory-bound mindset, the analytical version whose documented flaw day 49 closes.
- Day 42, Nsight Compute, where
dram__bytescomes from. - Day 43 and day 44, the matmul ladder climbing this chart rung by rung.
Sources
- Nsight Compute Profiling Guide, for the
dram__bytescounters and the profiler's own roofline charts: https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.