What are warp stall reasons in Nsight Compute?
The named categories Nsight Compute reports for why a warp could not issue, such as long scoreboard, barrier or MIO throttle.
Every cycle a resident warp does not issue an instruction, the hardware knows why, and Nsight Compute's Warp State Statistics section reports the tally by name. Long scoreboard means waiting on a global or local memory result. Barrier means parked at a __syncthreads(). LG throttle and MIO throttle mean the queues feeding the memory pipelines are full, so the warp cannot even hand its next memory instruction over. The profiling guide defines each category (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html , checked 2026-08-29).
The number they hang off is warp cycles per issued instruction: how many cycles an average warp spends between issues, split by reason. This is the diagnosis layer under latency hiding. High stall cycles are not themselves a problem; the machine is built to have warps waiting. They become the problem when the scheduler has no ready warp left, and the dominant reason tells you which fix applies: long scoreboard wants more independent loads or better locality, barrier wants less divergent work between syncs, throttle reasons want fewer, wider memory instructions rather than more parallelism.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), Nsight Compute 2024.3.2, captured 2026-09-01 with sudo ncu --set full --clock-control base. Day 42's two kernels, from the shipped profile/day42-naive and day42-tiled reports:
The naive kernel spends 15.6 of its 25.1 cycles between issues in LG throttle, 62.3 percent, saturated on the global-load path. The tiled version moved its loads to shared memory, and its top stall became MIO throttle at 15.4 of 28.7 cycles, 53.7 percent. Tiling moved the queue being waited on, and the kernel got 1.59x faster because the new queue drains faster.
| kernel | duration | dominant stall |
|---|---|---|
| naive matmul | 1.18 ms | LG throttle |
| tiled matmul | 744.64 us | MIO throttle |
The per-reason cycle counts come from the two reports' Warp State sections, which ship with the lesson so a reader without counter access can find the same rows.
Related terms
Where you meet this
- Day 42, reading an Nsight Compute report, the lesson that owns this term and ships the two reports.
- Day 45, occupancy, where the stall picture says whether more warps would help.
- Day 15, bank conflicts, a stall source with its own named counter.
- Colab setup, where the shipped reports stand in for counter access.
Sources
- Nsight Compute Profiling Guide, Warp Scheduler States, for the definition of each stall reason: https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.