What is kernel fusion in CUDA?
Merging several kernels into one so intermediate results stay in registers instead of round-tripping through global memory.
A chain of elementwise kernels, scale, activation, residual add, clamp, is four correct little programs and one bad memory access pattern. Each stage writes its result to global memory and the next stage reads it straight back, so the chain's traffic is dominated by intermediates no one asked to keep. Fuse the four bodies into one kernel and each value flows stage to stage through registers; the only global traffic left is the chain's real inputs and outputs. The speedup is not mysterious and it is not about instruction counts. It is bytes.
That is the number to compute before writing any code. Day 48's staged chain moves 36 bytes per element; the fused version moves 12, a 3.00x cut, and the measured speedup came out 3.08x. For a memory-bound chain, predicted bytes over bytes is the speedup, which also tells you when fusion is not worth it: if a stage is compute-bound, or the intermediate is genuinely needed elsewhere, the bytes do not shrink and neither does the time. Fusion also raises arithmetic intensity, since the same math now rides on fewer bytes, which is how a chain climbs the roofline without changing its arithmetic.
Fusion collapses launch overhead too, four launches into one, but on this chain at full size that was noise: the four launches cost tens of microseconds against milliseconds of traffic. Launch cost becomes the story only when the kernels are short, and a CUDA graph attacks that half without touching the bytes. The two optimizations are routinely conflated; they answer different problems.
The trade is register pressure and generality. A fused kernel holds more live values per thread, and a library call cannot fuse with your code around it, which is the standing argument for owning some kernels yourself.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75, captured 2026-09-01. Day 48 ran a scale-bias, GELU, residual-add, clamp chain over 16,777,216 floats, staged and fused, mean of 10 runs after 3 warm-ups:
| version | ms | bytes/elem | GB/s |
|---|---|---|---|
| staged chain, timed as one | 2.419 | 36 | 249.7 |
| chainFused | 0.785 | 12 | 256.5 |
staged / fused = 3.08, and the byte ratio is 3.00. Both versions run at the card's effective bandwidth (the copy floor measured 244.8 GB/s in the same program), so neither got faster at moving bytes; the fused one just moves a third as many. The Nsight Compute pass confirms the accounting from the DRAM side: dram__bytes for the fused kernel derives to 13.2 bytes per element against the 12 the arithmetic asks, with the individual staged kernels near 8.8 each for their 8.
Related terms
Where you meet this
- Day 48, kernel fusion, the lesson that owns this term and produced the numbers above.
- Day 39, Thrust, CUB and libcu++, where the inability of library calls to fuse is the case for writing kernels at all.
- Day 30, the memory-bound mindset, the bytes-first accounting this entry applies.
- Day 49, the measured roofline, where fewer bytes per FLOP moves a kernel along the intensity axis.
Sources
- Nsight Compute Profiling Guide, for the
dram__bytescounters used to confirm the byte accounting: https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.