What does a CUDA warp scheduler do?
The part of the SM that picks, every cycle, which resident warp gets to issue an instruction.
Nothing else on an SM decides what happens next. A CUDA core fetches nothing and holds no program counter; it takes an instruction and the operands to run it on. The guide gives the decision in one sentence: "At each instruction issue cycle, a warp scheduler selects a warp with threads ready to execute its next instruction (the active threads of the warp) and issues the instruction to those threads." Ready is the word carrying the weight. A warp waiting on a global load is not a candidate this cycle, and it costs nothing while it waits, because its registers and its program counter never left the chip.
That is the trade the whole machine is built on. A CPU core fights memory latency with cache and out-of-order execution and pays for both in silicon per core. An SM keeps many warps resident instead and switches between them at no cost: a T4 holds 32 resident warps per SM, and there is no context to save or restore, because all 32 sets of registers already sit in the register file. Latency hiding is not a technique you apply on top of this. It is what the scheduler does whenever you have handed it enough candidates, and occupancy is the size of that pool.
Which is why reading occupancy as throughput goes wrong. Once some warp is ready every cycle the scheduler is already issuing every cycle, and the extra warps you bought with a smaller register budget buy nothing back. The second trap is subtler: a split warp is still one candidate. When lanes disagree the scheduler has to issue the branch's two sides one after the other to the same warp, so the waste is invisible to any count of warps. Day 22 timed identical arithmetic at 3.272 ms when every warp split and 1.616 ms when the split fell on a warp boundary, which is warp divergence priced as issue slots rather than as branches.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75. Day 21 launches one block of 256 threads, has every thread write clock64() into its own slot of a device array, and prints each reading relative to the smallest one in that launch. clock64() counts per-multiprocessor cycles, so with one block in flight this is issue order and not a duration. Captured 2026-08-30 on the project's verification node.
| Warp | Cycles |
|---|---|
| 0 | 8 |
| 1 | 10 |
| 2 | 12 |
| 3 | 14 |
| 4 | 0 |
| 5 | 2 |
| 6 | 4 |
| 7 | 6 |
Eight warps, eight different readings, and warp 0 is not first. They reached that instruction two cycles apart in the order 4, 5, 6, 7, 0, 1, 2, 3. The scheduler picks, and it does not owe you your own index order.
A second pattern sits in the same column. Group the eight warps by index modulo four and each group holds two readings eight cycles apart: 4 and 0, 5 and 1, 6 and 2, 7 and 3. Turing puts four warp schedulers on an SM, so a pattern in fours is the shape you would expect. NVIDIA does not publish the assignment rule and this is one block on one card, so take it as a hint about where to look rather than a rule to code against.
The neighbouring column says what the scheduler's unit is. Inside every warp all 32 lanes reported the same cycle, and no two warps reported the same one.
Diagram
occupancy-stepper, preset issue-order: one SM's 32 warp slots with eight of them filled by a 256-thread block, each filled slot labelled with the cycle it issued on, and a step control that lights them in measured order rather than index order.
Alt text: "Eight warps of one block on an SM, each labelled with the cycle it reached the same instruction. They issue two cycles apart in the order four, five, six, seven, zero, one, two, three, so the scheduler's choice does not follow warp index."
Related terms
- streaming multiprocessor
- warp
- occupancy
- latency hiding
- warp stall reasons
- resident warps and blocks per SM
- CUDA core
- warp divergence
Where you meet this
- Day 2, how a GPU differs from a CPU, where the resident-warp budget the scheduler chooses from comes off the driver.
- Day 21, what is a warp, the lesson that owns this term and produced the table above.
- Day 22, warp divergence, for what one candidate warp costs when its lanes disagree.
- Day 42, Nsight Compute, where issue-slot utilization and the stall reasons say which cycles the scheduler wasted.
- Day 45, occupancy, where two configurations with different candidate pools reach the same speed.
- Colab setup, enough card to run day 21 and read your own issue order back, though not enough permission for the profiler counters.
Sources
- CUDA Programming Guide, "Hardware Multithreading" and "SIMT Architecture", for the warp scheduler's selection rule and the resident execution context: https://docs.nvidia.com/cuda/cuda-programming-guide/03-advanced/advanced-kernel-programming.html (checked 2026-08-30)
- Turing Tuning Guide 1.4.1.1, "Instruction Scheduling", for four warp schedulers per SM on this architecture: https://docs.nvidia.com/cuda/turing-tuning-guide/index.html (checked 2026-08-30)
- CUDA Programming Guide, compute capabilities appendix, Table 30, for 32 resident warps per SM at compute capability 7.5: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/compute-capabilities.html (checked 2026-08-30)
- CUDA C++ Programming Guide 7.13, the 12.6 archive, for
clock64()counting a per-multiprocessor cycle counter: https://docs.nvidia.com/cuda/archive/12.6.0/cuda-c-programming-guide/index.html (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.