← Glossary
CUDA glossaryExecution model
CC any

What is SIMT in CUDA?

Single instruction, multiple threads: the hardware issues one instruction to 32 threads that each keep their own registers and thread state. Since Volta, that state includes a per-thread program counter.

You write a kernel body as if one thread ran it. The SM "creates, manages, schedules, and executes threads in groups of 32 parallel threads called warps", and one instruction goes to all 32 lanes of a warp at once. Each lane keeps its own registers and its own index, so the code reads as scalar and executes 32 wide. That gap between how it reads and how it runs is where most beginner surprises live.

SIMD is the honest comparison, and it is not the same machine. On a vector unit you get one register that is 8 or 16 elements wide, and both the gathering and the branching are your problem. NVIDIA's own guide states the contrast: "Vector architectures, on the other hand, require the software to coalesce loads into vectors and manage divergence manually." SIMT moves both jobs into hardware. The memory system takes the warp's 32 addresses and merges them into the smallest set of transactions that covers them (memory coalescing), and a branch that splits a warp runs both sides with the wrong lanes masked off instead of refusing to compile.

Neither of those is free, which is the part a definition alone will not tell you. Because the warp is the unit of memory traffic, an address pattern the hardware cannot merge costs you real bandwidth: day 11 measured 232.9 GB/s when consecutive lanes read consecutive floats and 9.6 GB/s at stride 32, moving identical bytes. Because a divergent branch executes both paths, warp divergence charges you for work the lanes throw away. And since Volta each thread really does carry its own program counter, so "a warp moves in lockstep" is no longer safe to assume. That is independent thread scheduling, and it is why every warp primitive now takes a _sync suffix and an explicit mask.

Measured

On a Tesla T4 (driver 595.84, CUDA 12.6, built with nvcc -O3 -arch=sm_75), day 2 read warpSize off the driver: 32. The interesting number is what the SM does with a block that is not a multiple of it. A 48-thread block takes two warp slots and runs the second with 16 of its 32 lanes switched off for the block's whole life, and a 16-thread block takes a full slot to run 16 lanes. The allocation is ceil(T / 32) and it rounds up every time, so the SIMT width is not a suggestion you can round past. Captured 2026-08-30 on the project's verification node.

On a T4, 1 instruction covers 32 lanes with per-thread registers; when a branch diverges, each path executes with nonparticipating lanes masked until the warp reconverges.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.