← Glossary
CUDA glossaryTooling
CC 7.5

What is NVTX?

Annotations you add to your own code so a profiler timeline shows your phase names instead of anonymous bars.

A profiler knows what the driver knows: kernel names, API calls, copies. It does not know that three of those kernels together are "the residual check" and the next four are "one iteration". NVTX is how you tell it. Push a named range, do the work, pop it, and every profiler that speaks the protocol (Nsight Systems, Nsight Compute, the PyTorch profiler) nests your names over the raw events. The API is deliberately tiny: nvtxRangePushA("spmv"), nvtxRangePop(), header only, nothing to link (https://github.com/NVIDIA/NVTX , checked 2026-08-30).

The practical difference is between a timeline you scroll and a timeline you query. Day 41 profiles the same PageRank binary twice: nsys profile -t cuda gives a wall of cudaLaunchKernel rows, and -t cuda,nvtx gives ranges named scale, spmv, dangling-mass, combine, residual and copy-residual. With names in place, nsys stats --report nvtx_sum turns the recording into an answer table: total time and instance count per phase of your program, in your vocabulary.

The ranges cost almost nothing when no profiler is attached, so the habit that pays is annotating the program's permanent structure, not sprinkling ranges during a crisis and deleting them after. The phase that embarrassed day 41 (a four-byte device-to-host copy of the residual outweighing the entire SpMV by 13.7x) was only nameable because the range was already there.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), Nsight Systems 2024.3.2, captured 2026-09-01. From day 41's shipped day41-nvtx.nsys-rep, summarized with nsys stats --report nvtx_sum:

range total (ns) instances avg (ns)
iteration 53982410 359 150368.8
copy-residual 36972808 164 225444.0
spmv 2705884 359 7537.3
scale 3066298 359 8541.2

The copy-residual row is the finding: 36.97 ms total against spmv's 2.71 ms, and its average of 225444.0 ns for copying four bytes is the synchronization it drags, not the bytes. Without the ranges this is 359 anonymous launch rows; with them the loop's cost structure reads off one table.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.