What is Nsight Compute?
The kernel profiler, which replays one kernel and reports its memory chart, its stall reasons and how close it got to the hardware limits.
Where Nsight Systems shows you a timeline of everything, ncu picks one kernel launch, replays it as many times as its counters need, and answers the questions a timer cannot: how many transactions each warp's loads cost, what fraction of the hardware's limit the kernel reached, and why warps could not issue. The output is a .ncu-rep file you can open in the GUI or dump with --page details.
The first thing most people hit is not a metric, it is a permission error. On a stock Linux driver, plain ncu fails with ERR_NVGPUCTRPERM, because access to GPU performance counters is restricted to admin users; the restriction arrived with driver 418.43 on Linux (https://developer.nvidia.com/nvidia-development-tools-solutions-err_nvgpuctrperm-permission-issue-performance-counters , checked 2026-08-29). On the project's own node the cause is the driver default RmProfilingAdminOnly: 1, and sudo ncu works. On a hosted tier where you have no root, it does not, which is why every profiling lesson in this course ships its .ncu-rep files: reading a real report needs no GPU at all.
What the counters buy you is precision about mechanism. A timer told day 42 that tiling made the matmul 1.59x faster. The report says why, and also what tiling did not do: it did not cut load instructions per thread (1,088 against 1,024, it moved them to shared memory), and it barely changed DRAM traffic, because at this size the naive kernel's re-reads were already cache hits. The win lives in one pair of numbers, sectors moved, which no stopwatch can see.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), Nsight Compute 2024.3.2, captured 2026-09-01 under sudo ncu --set full --clock-control base. From the shipped profile/day42-naive and day42-tiled reports for a 512 cubed matmul:
| metric | naive | tiled |
|---|---|---|
| duration | 1.18 ms | 744.64 us |
| global load requests | 8388608 | 524288 |
| global load sectors | 16777077 | 2083878 |
| sectors per request | 2.0 | 4.0 (3.98) |
| L1/TEX hit rate | 87.58% | 5.13% |
| L2 hit rate | 97.93% | 97.05% |
| DRAM total | 4.74 MB | 4.64 MB |
The naive kernel wins on sectors per request and loses on sectors moved, 8.05x more, which is the memory chart's whole story. And both DRAM totals sit two orders of magnitude below what the naive kernel's arithmetic asks of global memory, which settles a question day 30's analytical roofline could only flag: the naive matmul's re-reads were served by cache, not DRAM.
Related terms
- Nsight Systems
- speed of light
- warp stall reasons
- memory transaction
- sector and cache line
- memory coalescing
Where you meet this
- Day 42, reading an Nsight Compute report, the lesson that owns this term and produced the table above.
- Day 11, memory coalescing, the first lesson whose model an ncu counter confirmed.
- Day 30, the memory-bound mindset, whose self-documented roofline flaw this report resolves.
- Day 49, the measured roofline, where
dram__bytes.sumreplaces counted-on-paper bytes. - Colab setup, where the counter permission is the reason the course ships its reports.
Sources
- NVIDIA, "ERR_NVGPUCTRPERM: Permission issue with Performance Counters", for the error and the driver versions that introduced the restriction: https://developer.nvidia.com/nvidia-development-tools-solutions-err_nvgpuctrperm-permission-issue-performance-counters (checked 2026-08-29)
- Nsight Compute Profiling Guide, for sectors, requests and the metric names quoted: https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.