What does compute-bound mean for a CUDA kernel?
A kernel whose math pipelines are the limit, so better memory access buys nothing.
The only compute-bound program on day 30 is the one that computes nothing anybody wants. It is eight chains of fused multiply-adds per thread with no memory in the loop, written to find the top of the card, and it reaches 8092.1 GFLOP/s. Everything else on that page, including a tiled matrix multiply, is waiting for data. That is not an accident of which kernels got picked. It is what the ratio between a modern GPU's two ceilings does to ordinary code.
Divide them and you get the ridge point, the arithmetic intensity at which the flat compute ceiling drops below the sloped bandwidth one. On this T4 it is 33.09 FLOP per byte, from 8092.1 GFLOP/s over a measured 244.5 GB/s. Right of that number your math pipelines are the limit, so coalescing, block shape and occupancy buy you nothing and the moves that pay are cheaper instructions, fewer of them, or a wider unit. Left of it none of that is true. Thirty-three flops for every byte crossing the bus is a demanding thing to ask of an algorithm, and almost nothing asks it by accident.
Look at the distance. Day 16's tiled matmul is the highest-intensity kernel in the first thirty days at 3.938 FLOP per byte, and it reaches 976.0 GFLOP/s against that 8092.1 ceiling while sitting on its bandwidth line. Tiling gives T / 4 FLOP per byte for a T by T float tile, so closing the remaining gap by tile size alone would need a tile many times wider than 16, and two float tiles that wide do not fit in the shared memory an SM has. The ladder therefore continues somewhere else: register tiling on days 43 and 44, which raises the reuse per byte without asking shared memory for more room, and half precision, which halves the denominator for the same tile.
Two things about the ridge point are worth carrying. It belongs to the card and the precision, not to your kernel, so quoting somebody else's is meaningless, and a card with a better flops-to-bandwidth ratio makes compute-bound harder to reach rather than easier. And it moves when the unit does: FP16 and tensor core paths raise the flat ceiling by a large factor, which pushes the ridge point right and can turn a kernel that was compute bound in FP32 back into a memory-bound one in FP16. That is also the honest limit of a FLOP roofline. It is blind to integer work, so a kernel that is genuinely limited by integer instructions reads as intensity zero and looks memory bound, and no amount of staring at the chart will say otherwise.
Measured
On a Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75, day 30 measures both ceilings itself rather than quoting a spec sheet, because a roofline drawn against a number nobody can reach makes every percentage on it wrong by the same unknown factor.
| Measured on the card | Value |
|---|---|
| copy, coalesced, in and out | 0.549 ms, 244.5 GB/s |
| fma, 8 chains, 2048 iterations | 1.327 ms, 8092.1 GFLOP/s |
| ridge point | 33.09 FLOP/byte |
| highest intensity in days 1 to 30 | 3.938 FLOP/byte, matmul tiled at 16 |
| that kernel's rate | 0.275 ms, 976.0 GFLOP/s |
Five of the five kernels day 30 places on this roofline sit left of 33.09, so not one of them is compute bound and a faster inner loop would move none of them. The crossover itself is not measured yet: the matmul ladder that is meant to reach it runs on days 43 and 44, and this entry gets the step where the verdict flips when those days run. Until then the number this page is worth reading for is 33.09, because it is a fact about the card that no rewrite changes.
One row from the same run must not be read as evidence of compute-bound behaviour. The naive matmul reports 615.9 GFLOP/s at 2466.1 GB/s, which is 1008.5 percent of the copy ceiling, and day 30's evidence file documents why before it prints the table: GB/s there comes from compulsory bytes while the kernel re-reads rows and columns out of L1 and L2. A kernel appearing to beat the memory system by a factor of ten is a counting artefact, never a discovery that it was compute bound all along.
Diagram
roofline-plotter, preset five-kernels: the flat compute ceiling drawn dashed at 8092.1 GFLOP/s, the sloped bandwidth ceiling solid, the ridge point marked at 33.09 FLOP per byte, and every kernel point far to its left. The region right of the ridge is empty, which is the picture.
Alt text: "A T4 roofline with the ridge point at 33.09 FLOP per byte. The flat compute ceiling sits at 8092.1 GFLOP/s and no kernel from the first thirty days is anywhere near it, because all of them are left of the ridge where bandwidth decides the speed."
Code
From code/day30-roofline/roofline.cu. This is what a compute-bound loop has to look like, and the shape is the lesson:
for (int k = 0; k < iters; ++k) {
#pragma unroll
for (int q = 0; q < kFmaChains; ++q) {
acc[q] = fmaf(acc[q], kFmaMul, kFmaAdd);
}
}
No load, no store, and eight independent accumulators so an FMA never waits on the one before it. iters is a run-time argument on purpose: a constexpr count would let nvcc unroll every multiply-add into straight-line code and measure the instruction cache instead. The program also recomputes one thread's chain on the host and compares, because a compiler that decided this loop was dead would hand back a ceiling the card cannot reach, and every percentage in the table is measured against it.
Related terms
Where you meet this
- Day 16, tiled matrix multiply, the highest-intensity kernel in the first thirty days and still a bandwidth-bound one.
- Day 30, arithmetic intensity and the memory-bound mindset, the lesson that owns this term.
- Day 43, optimizing matmul, where register tiling raises intensity past what a shared tile can reach.
- Day 49, building a roofline for your GPU, where you measure your own card's two ceilings and get your own ridge point.
- Learn CUDA without a GPU, because working out whether an algorithm could ever be compute bound needs no card.
Sources
- Nsight Compute Profiling Guide, section 2.9 "Roofline Charts", which names the flat part the peak performance boundary and the meeting point the ridge point: https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html (checked 2026-08-30)
- CUDA C++ Best Practices Guide, sections 9.2.1 and 9.2.2, on why the ceiling you compare against should be measured rather than quoted: https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html (checked 2026-08-30)
- Williams, Waterman and Patterson, "Roofline: An Insightful Visual Performance Model for Multicore Architectures", Communications of the ACM, April 2009: https://dl.acm.org/doi/10.1145/1498765.1498785 (checked 2026-08-30)
- CUDA Programming Guide, "Compute Capabilities", for the per-architecture shared memory and tensor core support that move this ceiling: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/compute-capabilities.html (checked 2026-08-29)
- "At what computational threshold can we say that CUDA is worth bringing into play?", the question this ratio answers: https://www.reddit.com/r/CUDA/comments/1gdvun8/cuda_vs_multithreading/ (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. This entry stays a draft until a named author and a different named reviewer sign it.