What is the PyTorch profiler?
PyTorch's operator-level profiler, which records CPU and CUDA activity, input shapes and traces that connect framework operations to device kernels.
torch.profiler.profile records the framework operations your program called and, when CUDA activity is enabled, the device work they launched. key_averages() turns those events into a table; record_shapes=True can split calls with the same operator name by input shape, and export_chrome_trace() writes a timeline for a trace viewer.
The table has inclusive and exclusive columns. Total time includes children, so adding totals double-counts nested dispatch. Self time is the portion attributed to that row rather than its children, but it is still an attribution produced by the installed profiler, not a promise that only raw kernel-name rows will be nonzero. On the verified torch 2.13.0 run, the same GEMM activity appeared under aten::addmm, Unrecognized and volta_sgemm_128x128_tn. Read the call tree and trace before treating rows as independent work.
The profiler uses CUPTI for CUDA activity. That matters when another profiler wraps the process. Nsight Systems also subscribes to CUPTI, and day 88's nsys-wrapped rerun made torch profiler warn that CUPTI already had a subscriber. The correct evidence split is:
- the clean standalone run supplies the torch profiler CUDA tables and Chrome traces;
- the wrapped run supplies the valid Nsight Systems CUDA and NVTX reports;
- no torch CUDA table from the subscriber-collision run is used.
This is a collection conflict, not a requirement for root or GPU performance-counter permission. Both clean activity-tracing paths worked unprivileged on the verification node.
Measured
On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 88 profiled a Linear followed by either five eager elementwise kernels or one fused operator. In the clean torch profiler table, the eager elementwise rows summed to 1.5118 ms and the fused kernel took 0.4104 ms, a 3.68x improvement against an 11-to-3 byte ratio. Separate CUDA-event timing measured the complete blocks at 7.7794 ms and 6.6554 ms, only 1.1689x apart because the unchanged GEMM still dominated.
The run exported both Chrome traces. A separate Nsight Systems capture ranked volta_sgemm_128x128_tn first by a wide margin and retained its NVTX and CUDA-kernel CSV summaries. Its NVTX PushPop durations are host range durations around asynchronous launches, not replacements for the CUDA-event block timings.
Related terms
Where you meet this
- Day 88, profiling a PyTorch model, which produced the tables and traces above.
- Day 87, custom PyTorch operators, where a kernel becomes an operator the profiler can name.
- Day 41, Nsight Systems, for the process-wide timeline used beside the clean framework profile.
Sources
- PyTorch profiler API, including activities, shape recording, key averages and trace export: https://docs.pytorch.org/docs/stable/profiler.html (checked 2026-09-01)
- PyTorch profiler recipe, including CUDA kernel rows and interpreting profiler tables: https://docs.pytorch.org/tutorials/recipes/recipes/profiler_recipe.html (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.