← Glossary
CUDA glossaryTooling
CC 7.5

What is Nsight Systems?

The timeline profiler, which shows what the CPU, the copies and the kernels were each doing and when.

nsys profile ./app records everything that happened, runtime calls on the CPU side, kernels and copies on the GPU side, and lays it out against one clock. That is a different question from the one Nsight Compute answers. ncu replays a single kernel and explains its inner life; nsys watches the whole program once and shows where the time went between the kernels: gaps, synchronizations, copies, and phases that should overlap but do not. When a program is slower than the sum of its kernels, the difference lives on this timeline and nowhere else.

The tool is only as readable as the program is labeled, which is what NVTX ranges are for: wrap each phase in a named range and the timeline shows spmv and copy-residual instead of anonymous bars. You do not need the GUI to read the result either. nsys stats --report nvtx_sum report.nsys-rep prints a per-range summary table, and every report name is documented in the analysis guide (https://docs.nvidia.com/nsight-systems/AnalysisGuide/index.html , checked 2026-08-30).

Unlike the counter-based profiler, nsys did not need root on the project's node, and its overhead model is tracing rather than replay, so it is the tool you reach for first. The course's rule of thumb: nsys to find which phase is wrong, ncu to find out why that phase's kernel is slow.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), Nsight Systems 2024.3.2, captured 2026-09-01. Day 41 profiled its PageRank loop with every phase inside an NVTX range and summarized the shipped report with nsys stats --report nvtx_sum:

range total (ns) instances
copy-residual 36972808 164
spmv 2705884 359
dangling-mass 5125947 359
combine 2687104 359

The gap the timeline revealed: the loop spent 36.97 ms copying a four-byte residual back to the host against 2.71 ms doing the actual sparse matrix work, 13.7x, because each copy drags a synchronization with it. The same run priced the design question: checking convergence every step costs 1.14x never checking. A timer around the whole loop shows the 1.14x; only the timeline says which bar to blame.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.