What is a warp in CUDA?
The 32 threads an SM issues together, sharing one instruction stream and one set of memory requests.
Nothing in your source names a warp. You pick a block size in the launch configuration, and the SM cuts that block into groups of 32 by thread ID: "The way a block is partitioned into warps is always the same; each warp contains threads of consecutive, increasing thread IDs with the first warp containing thread 0." Threads 0 to 31 are warp 0 on every card and every run. After that the warp is the thing the hardware bills. One instruction goes to all 32 lanes, one branch decision covers all 32, and 32 addresses arrive at the memory pipe as one request.
The division rounds up, which is where the first surprise sits. A block of T threads occupies ceil(T / 32) warps, so 33 threads take two slots and the second holds one live lane. A T4 SM has 32 resident warp slots and spends two of them on that block, exactly as many as it spends on a block of 64. Occupancy cannot separate the two, because it counts warps rather than working lanes, so 31 dead lanes never reach the metric people check.
Then there is the number itself, which the documentation states twice at two different strengths. warpSize is "A run-time value defined as the number of threads in a warp, commonly 32". The hardware multithreading section that defines ceil(T / Wsize) says "Wsize is the warp size, which is equal to 32". Both readings have a practical edge: warpSize is a run-time value, so it cannot size a __shared__ array or appear in a static_assert, and 32 hardcoded everywhere is an assumption nobody rechecks. Declare constexpr int kWarpSize = 32; for the compile-time uses, read warpSize once, and compare them.
Thinking in warps pays because the expensive mistakes are warp-shaped. Addresses that consecutive lanes do not share turn one request into many, which is memory coalescing. A predicate that changes inside a run of 32 makes the warp execute both sides, which is warp divergence: day 22 timed threadIdx.x & 1 at 3.272 ms against 1.616 ms for the same arithmetic split on a warp boundary, 2.02x paid for where the split fell and nothing else.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75. Day 21 runs one block at a time and has every thread write clock64() and __activemask() into its own slot of a device array, so nothing is inferred from print order. Captured 2026-08-30 on the project's verification node.
| Threads in the block | Warps | Lanes in the last warp | Its active mask | Distinct clock64 values |
|---|---|---|---|---|
| 32 | 1 | 32 | 0xffffffff |
1 |
| 33 | 2 | 1 | 0x00000001 |
2 |
| 48 | 2 | 16 | 0x0000ffff |
2 |
| 64 | 2 | 32 | 0xffffffff |
2 |
| 256 | 8 | 32 | 0xffffffff |
8 |
The 33-thread row is the whole term in one line. One thread past a full warp costs a second warp slot, and the mask says so in hardware: 0x00000001, one bit, 31 lanes allocated and idle for as long as the block lives.
The right-hand column is the other half of the definition. Inside a warp every lane reported the same clock64() value, and no two warps reported the same one, so the 256-thread block came back with eight warps and eight numbers. That is what "together" means, measured rather than asserted. Two caveats travel with it: clock64() reads a per-multiprocessor counter, which is why every launch here is a single block, and a T4 is compute capability 7.5, so independent thread scheduling is on and one instruction's agreement is an observation rather than a guarantee.
Diagram
occupancy-stepper, preset partial-warp: one SM's 32 warp slots down the side, a block of 48 threads filling two of them, the first slot 32 lanes lit and the second 16 lit beside 16 grey.
Alt text: "A 48-thread block takes two of an SM's 32 warp slots. The first warp runs 32 lanes, the second runs 16 alongside 16 that were never launched and stay idle for the life of the block."
Code
From code/day21-warps/warps.cu. Both functions are constexpr, so the counts on this page are checked by the compiler and the file stops building if they stop being true.
// Warps a block of t threads occupies. This is the guide's own formula,
// ceil(T / Wsize), and it rounds up: a block of 48 threads occupies two warps
// and the second one carries 16 lanes that were never launched.
constexpr int warpsInBlock(int t) {
return (t + kWarpSize - 1) / kWarpSize;
}
// Threads that actually exist in warp `w` of a block of `t` threads. Every
// warp but the last holds 32; the last holds the remainder.
constexpr int lanesInWarp(int t, int w) {
return (t - w * kWarpSize < kWarpSize) ? (t - w * kWarpSize) : kWarpSize;
}
Related terms
- lane
- SIMT
- warp divergence
- warp scheduler
- warp shuffle
- streaming multiprocessor
- resident warps and blocks per SM
- memory coalescing
Where you meet this
- Day 2, how a GPU differs from a CPU, where the driver first reports a warp size of 32.
- Day 10, threads per block, the block-size sweep this rounding rule decides.
- Day 21, what is a warp, the lesson that owns this term and produced the table above.
- Day 22, warp divergence, for what a warp costs when its lanes disagree.
- Day 23, warp shuffles, for the instructions that address lanes on purpose.
invalid configuration argument, the launch failure when a block size is illegal rather than merely wasteful.
Sources
- CUDA Programming Guide, "SIMT Architecture" and "Hardware Multithreading", for the partitioning rule,
ceil(T / Wsize)and the group-of-32 wording: https://docs.nvidia.com/cuda/cuda-programming-guide/03-advanced/advanced-kernel-programming.html (checked 2026-08-30) - CUDA Programming Guide, C++ language extensions, for
warpSizeas "A run-time value" and for__activemask(): https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/cpp-language-extensions.html (checked 2026-08-30) - CUDA Programming Guide, compute capabilities appendix, Table 30, for 32 resident warps per SM at compute capability 7.5: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/compute-capabilities.html (checked 2026-08-30)
- CUDA C++ Programming Guide 7.13, the 12.6 archive, for what
clock64()counts: https://docs.nvidia.com/cuda/archive/12.6.0/cuda-c-programming-guide/index.html (checked 2026-08-30) - "How do CUDA blocks/warps/threads map onto CUDA cores?", 86,520 views and the most-read concept question in the tag: https://stackoverflow.com/questions/10460742/how-do-cuda-blocks-warps-threads-map-onto-cuda-cores (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.