← Glossary
CUDA glossaryPerformance
CC 7.5

What is kernel fusion in CUDA?

Merging several kernels into one so intermediate results stay in registers instead of round-tripping through global memory.

A chain of elementwise kernels, scale, activation, residual add, clamp, is four correct little programs and one bad memory access pattern. Each stage writes its result to global memory and the next stage reads it straight back, so the chain's traffic is dominated by intermediates no one asked to keep. Fuse the four bodies into one kernel and each value flows stage to stage through registers; the only global traffic left is the chain's real inputs and outputs. The speedup is not mysterious and it is not about instruction counts. It is bytes.

That is the number to compute before writing any code. Day 48's staged chain moves 36 bytes per element; the fused version moves 12, a 3.00x cut, and the measured speedup came out 3.08x. For a memory-bound chain, predicted bytes over bytes is the speedup, which also tells you when fusion is not worth it: if a stage is compute-bound, or the intermediate is genuinely needed elsewhere, the bytes do not shrink and neither does the time. Fusion also raises arithmetic intensity, since the same math now rides on fewer bytes, which is how a chain climbs the roofline without changing its arithmetic.

Fusion collapses launch overhead too, four launches into one, but on this chain at full size that was noise: the four launches cost tens of microseconds against milliseconds of traffic. Launch cost becomes the story only when the kernels are short, and a CUDA graph attacks that half without touching the bytes. The two optimizations are routinely conflated; they answer different problems.

The trade is register pressure and generality. A fused kernel holds more live values per thread, and a library call cannot fuse with your code around it, which is the standing argument for owning some kernels yourself.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75, captured 2026-09-01. Day 48 ran a scale-bias, GELU, residual-add, clamp chain over 16,777,216 floats, staged and fused, mean of 10 runs after 3 warm-ups:

version ms bytes/elem GB/s
staged chain, timed as one 2.419 36 249.7
chainFused 0.785 12 256.5

staged / fused = 3.08, and the byte ratio is 3.00. Both versions run at the card's effective bandwidth (the copy floor measured 244.8 GB/s in the same program), so neither got faster at moving bytes; the fused one just moves a third as many. The Nsight Compute pass confirms the accounting from the DRAM side: dram__bytes for the fused kernel derives to 13.2 bytes per element against the 12 the arithmetic asks, with the individual staged kernels near 8.8 each for their 8.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.