What is Nsight Systems?
The timeline profiler, which shows what the CPU, the copies and the kernels were each doing and when.
nsys profile ./app records everything that happened, runtime calls on the CPU side, kernels and copies on the GPU side, and lays it out against one clock. That is a different question from the one Nsight Compute answers. ncu replays a single kernel and explains its inner life; nsys watches the whole program once and shows where the time went between the kernels: gaps, synchronizations, copies, and phases that should overlap but do not. When a program is slower than the sum of its kernels, the difference lives on this timeline and nowhere else.
The tool is only as readable as the program is labeled, which is what NVTX ranges are for: wrap each phase in a named range and the timeline shows spmv and copy-residual instead of anonymous bars. You do not need the GUI to read the result either. nsys stats --report nvtx_sum report.nsys-rep prints a per-range summary table, and every report name is documented in the analysis guide (https://docs.nvidia.com/nsight-systems/AnalysisGuide/index.html , checked 2026-08-30).
Unlike the counter-based profiler, nsys did not need root on the project's node, and its overhead model is tracing rather than replay, so it is the tool you reach for first. The course's rule of thumb: nsys to find which phase is wrong, ncu to find out why that phase's kernel is slow.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), Nsight Systems 2024.3.2, captured 2026-09-01. Day 41 profiled its PageRank loop with every phase inside an NVTX range and summarized the shipped report with nsys stats --report nvtx_sum:
| range | total (ns) | instances |
|---|---|---|
| copy-residual | 36972808 | 164 |
| spmv | 2705884 | 359 |
| dangling-mass | 5125947 | 359 |
| combine | 2687104 | 359 |
The gap the timeline revealed: the loop spent 36.97 ms copying a four-byte residual back to the host against 2.71 ms doing the actual sparse matrix work, 13.7x, because each copy drags a synchronization with it. The same run priced the design question: checking convergence every step costs 1.14x never checking. A timer around the whole loop shows the 1.14x; only the timeline says which bar to blame.
Related terms
Where you meet this
- Day 41, Nsight Systems, the lesson that owns this term and ships the report above.
- Day 42, Nsight Compute, the kernel-level half of the same toolbox.
- Day 9, why your GPU code looks slower than your CPU, the copy-dominated timeline every beginner records first.
- Colab setup, where reading the shipped
.nsys-repreplaces recording your own.
Sources
- Nsight Systems User Guide, "Profiling from the CLI": https://docs.nvidia.com/nsight-systems/UserGuide/index.html (checked 2026-08-30)
- Nsight Systems Post-Collection Analysis Guide, for
nsys statsreport names includingnvtx_sum: https://docs.nvidia.com/nsight-systems/AnalysisGuide/index.html (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.