What is arithmetic intensity in CUDA?
Flops per byte moved, the one number that says whether a kernel is limited by math or by memory.
Both halves of the fraction are yours to count, and nothing checks either one. That is the whole character of this number. A profiler will hand you achieved GFLOP/s and achieved GB/s, but the ratio that decides which ceiling applies to your kernel comes out of reading your own source and counting, which means it is available before you own a GPU and it is wrong the moment your count is. Day 30 works four kernels by hand on paper and then has the program print its counts beside its times, so a bad count shows up on the page rather than folding silently into a ratio.
Count the numerator wrong in two directions. A fused multiply-add is two flops, not one, and every convention that quotes GFLOP/s counts it that way, so halving your numerator can move a kernel across the ridge point on paper. In the other direction, source that looks heavy often is not: sqrtf and expf are a handful of SASS instructions, and integer arithmetic is worth zero on a FLOP roofline no matter how much of it there is. Day 30's grayscale kernel does three integer multiplies, two adds and a shift, and its intensity is 0.000, the same as a transpose that does no arithmetic at all.
The denominator is the subtler half. Bytes here means the traffic the algorithm has to ask global memory for, not the traffic DRAM sees, because the caches serve part of the request without telling you. This is what makes intensity something you can change: tiling a matmul does not touch the arithmetic at all, it removes repeat reads, and a T by T float tile takes the ratio from 2K / 8K to T / 4 FLOP per byte. At day 16's tile of 16 that is a factor of sixteen out of a change no arithmetic sees. It is also what lets a measured point sit above the ceiling the model puts it under, and that gap is a measurement of the reuse you were getting for free rather than a broken chart.
Then the part people carry away backwards: this is a diagnosis, not a score. Day 20's separable blur computes the same result in two 5-tap passes instead of one 25-tap pass, so its numerator falls by more than half with the traffic unchanged, its intensity goes down, and it is faster. Raising intensity helps only when you raised it by moving fewer bytes. Lowering it helps whenever you did less work for the same traffic. What the number is for is picking which optimisation can possibly pay: coalescing, block size and occupancy move a kernel up towards the ceiling it is already under, and only tiling, fusion and keeping data resident move it right to a different one.
Measured
On a Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75, day 30 counts five kernels from its own source, runs all five in one program, and measures both of the card's ceilings in the same run: a coalesced copy at 244.5 GB/s and a fused multiply-add chain at 8092.1 GFLOP/s, giving a ridge point of 33.09 FLOP per byte.
| Kernel | FLOP/byte | Time | GFLOP/s |
|---|---|---|---|
| vector add | 0.083 | 0.786 ms | 21.3 |
| transpose, naive | 0.000 | 1.302 ms | 0.0 |
| block reduction | 0.248 | 0.563 ms | 29.7 |
| matmul, naive | 0.250 | 0.436 ms | 615.9 |
| matmul, tiled 16 | 3.938 | 0.275 ms | 976.0 |
Two rows next to each other carry the point. The naive matmul and the block reduction sit at 0.250 and 0.248, a kernel that looks a hundred times heavier landing on the same ratio as one that adds up an array, because both read roughly four bytes per operation. Tiling the same matmul moves it to 3.938 with no change to its arithmetic, and even that is two orders of magnitude short of 33.09, which is why all five are bandwidth bound.
One caveat travels with that table and day 30's evidence file states it before the numbers. The same run reports the naive matmul at 2466.1 GB/s, which is 1008.5 percent of this card's measured copy ceiling. That is the counting rule above showing its edge: GB/s there is computed from compulsory bytes while the kernel re-reads the same rows and columns out of L1 and L2, so the figure describes traffic that never reached DRAM and must not be read as achieved bandwidth. The tiled row at 101.3 percent is the same effect nearly gone, because tiling is what removed the repeats. Separating the two needs dram__bytes from a profiler, which is day 42.
Diagram
roofline-plotter, preset five-kernels: a log-log chart with intensity on the x axis, the sloped bandwidth ceiling and the flat compute ceiling drawn as solid and dashed lines, the ridge point marked and labelled 33.09 FLOP per byte, and one point per kernel. Walking the points reads each verdict aloud.
Alt text: "Five kernels on a T4 roofline. All five sit left of a ridge point at 33.09 FLOP per byte, so the sloped bandwidth ceiling is the one that binds them, and the tiled matmul at 3.938 is the furthest right by a factor of fifteen."
Code
From code/day30-roofline/roofline.cu. The counts are literals worked from each kernel's source and nowhere else, so they sit next to the times on the page rather than inside a ratio:
const double flops[kNumRows] = {n, 0.0, n - blocks, 2.0 * d3, 2.0 * d3};
const double moved[kNumRows] = {
3.0 * n * sizeof(float), 2.0 * n * sizeof(float),
(n + blocks) * sizeof(float), (2.0 * d3 + d2) * sizeof(float),
(2.0 * d3 / kTileDim + d2) * sizeof(float)};
Read the last two entries together. Both matmuls do 2 * d3 flops, so the numerator is identical, and the only difference in the whole table is 2.0 * d3 / kTileDim instead of 2.0 * d3 in the denominator. That division by the tile dimension is the entire gain from tiling, expressed the way the model sees it.
Related terms
Where you meet this
- Day 12, why transposing a matrix is slow, the kernel whose intensity is zero and stays zero.
- Day 16, tiled matrix multiply, where the tile dimension turns into the denominator.
- Day 30, arithmetic intensity and the memory-bound mindset, the lesson that owns this term.
- Day 49, building a roofline for your GPU, where you measure both ceilings on your own card and plot ten kernels.
- Learn CUDA without a GPU, because every ratio on this page is arithmetic you can do on a laptop with no card in it.
Sources
- Nsight Compute Profiling Guide, section 2.9 "Roofline Charts", for the boundaries, the ridge point and the FLOP-per-byte unit: https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html (checked 2026-08-30)
- CUDA C++ Best Practices Guide, sections 9.2.1 and 9.2.2, theoretical against effective bandwidth, which is the formula day 30's copy ceiling uses: https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html (checked 2026-08-30)
- Williams, Waterman and Patterson, "Roofline: An Insightful Visual Performance Model for Multicore Architectures", Communications of the ACM, April 2009: https://dl.acm.org/doi/10.1145/1498765.1498785 (checked 2026-08-30)
- The learner who optimised the arithmetic of a memory-bound kernel for a week: https://www.reddit.com/r/CUDA/comments/1qkp3lb/my_first_optimization_lesson_was_stop_guessing_lol/ (checked 2026-08-29)
- Programming Massively Parallel Processors, 4th edition, chapter 5, on the compute to global memory access ratio: https://shop.elsevier.com/books/programming-massively-parallel-processors/hwu/978-0-323-91231-0 (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. This entry stays a draft until a named author and a different named reviewer sign it.