← Glossary
CUDA glossaryPerformance
CC 7.5

What is loop unrolling in CUDA?

The compiler replacing a loop with repeated copies of its body, which removes branch work and exposes independent instructions.

A loop pays per iteration: the counter update, the compare, the branch. Unroll it and those disappear, but the payment was never the main point on a GPU. The real prize is what unrolling makes visible: with the bodies laid out straight, the compiler can see which iterations are independent, schedule their loads early, and keep more results in flight per thread, which is latency hiding bought at compile time. The price is registers, because everything scheduled early needs somewhere to live. #pragma unroll requests it per loop; the compiler also unrolls on its own whenever the trip count is knowable.

That last clause is the part people miss, and day 46 caught it in the disassembly: the "rolled" version of its 64-tap fold shows 29 FFMAs inside its loop body. The compiler had already partially unrolled it, keeping a small loop around a widened body, without being asked. So "my loop" and "the loop the machine runs" diverge silently, and only the SASS back-edge count says which you got.

The other lesson in the same table is that unrolling is not free speed. The fully unrolled version dropped both back edges and cost 20 more registers per thread, and the time barely moved, 0.113 against 0.114 ms, because this kernel's limit was memory, not branch overhead. Unrolling helps when instruction issue or dependent-chain latency is the wall; when bytes are the wall, it spends registers on nothing. The measured occupancy cost of a register budget is day 17's topic, and an unroll pragma is one of the easiest ways to grow that budget by accident.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75, captured 2026-09-01. Day 46's same fold, two builds, shapes counted in the disassembly:

kernel back edges FFMA regs ms
decayRolled 2 29 42 0.114
decayUnrolled 0 63 62 0.113

The two kernels' outputs differ at 0 of 1,049,187 elements, so this is one program compiled twice, and the table is the whole trade laid bare: the unroll removed every branch, doubled the visible FFMAs, charged 20 registers, and bought one microsecond per thousand. The 63 rather than 64 FFMAs is its own small find: the first tap multiplies an accumulator still holding zero, so the compiler folded the dead multiply into an FADD.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.