← Glossary
CUDA glossaryPerformance
CC any

What does memory-bound mean for a CUDA kernel?

A kernel that spends its time waiting for data, so making the math faster changes nothing.

The phrase names which ceiling applies to you. It says nothing at all about how close you are to that ceiling, and reading it as one claim instead of two is where the wasted week comes from. Somebody on r/CUDA spent one: "I assumed the kernel was compute-bound because the mathematical operations seemed complex ... the problem was actually quite simple: it was memory-bound due to non-coalesced global memory accesses, so those fancy changes were completely useless." The verdict was available from one division before any of that work started.

The division is arithmetic intensity against the card's ridge point, which is its peak flops divided by its peak bytes per second. Day 30 measured both ends on a T4 and got 33.09 FLOP per byte. Below that number the sloped ceiling is lower than the flat one, so bandwidth decides your speed no matter what the arithmetic looks like. Every kernel in the first thirty days of this course sits left of it, most of them by two orders of magnitude, which is worth saying plainly: memory bound is the normal state of GPU code, not a diagnosis you have to earn. The interesting question is never whether you are memory bound, it is what to do about it.

That is where the second number matters. A naive transpose runs at 42.2 percent of this card's copy ceiling and vector add runs at 104.8 percent, and both are memory bound. The advice is opposite. The transpose has most of its ceiling unclaimed, so the job is coalescing, block shape and occupancy, all of which move a kernel up towards a ceiling it is already under. Vector add has nothing left under the sloped line, so nothing about how it touches memory can help it and the only remaining move is to make it ask for fewer bytes, by tiling, by fusing it with whatever runs next, or by not moving the data at all. One verdict, two completely different afternoons.

A point can also read above 100 percent, and there are two reasons that are not measurement errors. The mild one is the ceiling itself: a copy ceiling is half reads and half writes, so a kernel with a different mix can beat it slightly, which is vector add at 104.8. The loud one is the byte count. Intensity is computed from compulsory traffic, the minimum the algorithm must ask for, so a kernel with reuse moves less than that and the caches absorb the difference. Day 30's naive matmul reports 1008.5 percent of the copy ceiling for exactly this reason, and its evidence file flags the row before anyone can quote it: that figure is compulsory bytes over elapsed time, not achieved bandwidth, and nothing on the card moved 2466.1 GB/s. Treat any percentage above 100 as a message about your cache hit rate.

Measured

On a Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75, day 30 measures both ceilings in the same program that runs the kernels: a coalesced copy at 244.5 GB/s in 0.549 ms, and a fused multiply-add chain at 8092.1 GFLOP/s in 1.327 ms. Their ratio is the ridge point, 33.09 FLOP per byte. No vendor figure appears anywhere in it.

Kernel FLOP/byte Time GB/s Share of the ceiling
vector add 0.083 0.786 ms 256.2 104.8%
transpose, naive 0.000 1.302 ms 103.1 42.2%
block reduction 0.248 0.563 ms 119.7 48.9%
matmul, tiled 16 3.938 0.275 ms 247.8 101.3%

Day 30 placed five kernels and all five sat left of 33.09. The one missing from the table is the naive matmul at 0.250 FLOP per byte, left out on purpose: its 2466.1 GB/s is the compulsory-bytes artefact described above, documented at the end of day 30's evidence file, and it belongs in a discussion of counting rather than in a list of achieved rates.

The rest of the column is the useful part. Vector add and the tiled matmul are done, sitting on their sloped ceiling with a percent or two of slack from the copy ceiling's read-write mix. The transpose and the block reduction are not, at 42.2 and 48.9 percent, and both of those gaps have names the earlier days already gave them: a strided write in one, and a reduction that only reads one element per thread in the other. The roofline says which ceiling. It cannot say why you are at 42 percent of it, and day 42 is where the profiler answers that.

Diagram

roofline-plotter, preset five-kernels: a log-log chart with the sloped bandwidth ceiling and the flat compute ceiling, the ridge point marked at 33.09 FLOP per byte, and every kernel point drawn at its measured rate rather than on the line, so the vertical gap between a point and the ceiling above it is the headroom.

Alt text: "Four kernels under a T4's sloped bandwidth ceiling. Vector add and the tiled matmul sit on the line with nothing left to gain from memory work, while the naive transpose sits at forty-two percent of it, which is headroom a better access pattern can still claim."

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. This entry stays a draft until a named author and a different named reviewer sign it.