← Glossary
CUDA glossaryTooling
CC 7.5

What is Nsight Compute?

The kernel profiler, which replays one kernel and reports its memory chart, its stall reasons and how close it got to the hardware limits.

Where Nsight Systems shows you a timeline of everything, ncu picks one kernel launch, replays it as many times as its counters need, and answers the questions a timer cannot: how many transactions each warp's loads cost, what fraction of the hardware's limit the kernel reached, and why warps could not issue. The output is a .ncu-rep file you can open in the GUI or dump with --page details.

The first thing most people hit is not a metric, it is a permission error. On a stock Linux driver, plain ncu fails with ERR_NVGPUCTRPERM, because access to GPU performance counters is restricted to admin users; the restriction arrived with driver 418.43 on Linux (https://developer.nvidia.com/nvidia-development-tools-solutions-err_nvgpuctrperm-permission-issue-performance-counters , checked 2026-08-29). On the project's own node the cause is the driver default RmProfilingAdminOnly: 1, and sudo ncu works. On a hosted tier where you have no root, it does not, which is why every profiling lesson in this course ships its .ncu-rep files: reading a real report needs no GPU at all.

What the counters buy you is precision about mechanism. A timer told day 42 that tiling made the matmul 1.59x faster. The report says why, and also what tiling did not do: it did not cut load instructions per thread (1,088 against 1,024, it moved them to shared memory), and it barely changed DRAM traffic, because at this size the naive kernel's re-reads were already cache hits. The win lives in one pair of numbers, sectors moved, which no stopwatch can see.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), Nsight Compute 2024.3.2, captured 2026-09-01 under sudo ncu --set full --clock-control base. From the shipped profile/day42-naive and day42-tiled reports for a 512 cubed matmul:

metric naive tiled
duration 1.18 ms 744.64 us
global load requests 8388608 524288
global load sectors 16777077 2083878
sectors per request 2.0 4.0 (3.98)
L1/TEX hit rate 87.58% 5.13%
L2 hit rate 97.93% 97.05%
DRAM total 4.74 MB 4.64 MB

The naive kernel wins on sectors per request and loses on sectors moved, 8.05x more, which is the memory chart's whole story. And both DRAM totals sit two orders of magnitude below what the naive kernel's arithmetic asks of global memory, which settles a question day 30's analytical roofline could only flag: the naive matmul's re-reads were served by cache, not DRAM.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.