What is NVTX?
Annotations you add to your own code so a profiler timeline shows your phase names instead of anonymous bars.
A profiler knows what the driver knows: kernel names, API calls, copies. It does not know that three of those kernels together are "the residual check" and the next four are "one iteration". NVTX is how you tell it. Push a named range, do the work, pop it, and every profiler that speaks the protocol (Nsight Systems, Nsight Compute, the PyTorch profiler) nests your names over the raw events. The API is deliberately tiny: nvtxRangePushA("spmv"), nvtxRangePop(), header only, nothing to link (https://github.com/NVIDIA/NVTX , checked 2026-08-30).
The practical difference is between a timeline you scroll and a timeline you query. Day 41 profiles the same PageRank binary twice: nsys profile -t cuda gives a wall of cudaLaunchKernel rows, and -t cuda,nvtx gives ranges named scale, spmv, dangling-mass, combine, residual and copy-residual. With names in place, nsys stats --report nvtx_sum turns the recording into an answer table: total time and instance count per phase of your program, in your vocabulary.
The ranges cost almost nothing when no profiler is attached, so the habit that pays is annotating the program's permanent structure, not sprinkling ranges during a crisis and deleting them after. The phase that embarrassed day 41 (a four-byte device-to-host copy of the residual outweighing the entire SpMV by 13.7x) was only nameable because the range was already there.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), Nsight Systems 2024.3.2, captured 2026-09-01. From day 41's shipped day41-nvtx.nsys-rep, summarized with nsys stats --report nvtx_sum:
| range | total (ns) | instances | avg (ns) |
|---|---|---|---|
| iteration | 53982410 | 359 | 150368.8 |
| copy-residual | 36972808 | 164 | 225444.0 |
| spmv | 2705884 | 359 | 7537.3 |
| scale | 3066298 | 359 | 8541.2 |
The copy-residual row is the finding: 36.97 ms total against spmv's 2.71 ms, and its average of 225444.0 ns for copying four bytes is the synchronization it drags, not the bytes. Without the ranges this is 359 anonymous launch rows; with them the loop's cost structure reads off one table.
Related terms
Where you meet this
- Day 41, Nsight Systems, the lesson that owns this term and produced both traces.
- Day 40, the PageRank capstone, the program the ranges are wrapped around.
- Day 9, timing with events, the other way this course names where time goes.
Sources
- NVIDIA/NVTX, the v3 headers, where
nvtxRangePushAandnvtxRangePopare declared and the header-only rule is stated: https://github.com/NVIDIA/NVTX (checked 2026-08-30) - Nsight Systems Post-Collection Analysis Guide, for the
nvtx_sumreport: https://docs.nvidia.com/nsight-systems/AnalysisGuide/index.html (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.