← Glossary
CUDA glossaryPerformance
CC 7.5

What is speed of light in Nsight Compute?

Nsight Compute's headline percentage: how close a kernel gets to the hardware's maximum for compute and for memory.

It is the first section of every Nsight Compute report, and it answers the triage question before any detail: of what this GPU could theoretically sustain, what fraction did this kernel achieve, once for the SM's compute pipelines and once for the memory system. The name is the physics joke made policy; the guide defines the metrics as achieved percentages of a theoretical maximum (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html , checked 2026-08-29). Read it like a roofline collapsed to two numbers: whichever percentage is higher is the wall you are closer to, and if both are low, the kernel's problem is latency or launch shape rather than any throughput limit.

The subtlety is that "memory" is a hierarchy, and the section reports it at more than one level. A kernel can sit at 60 percent of memory throughput while its DRAM row reads 1 percent, which means the traffic is being absorbed by L1 and L2 and hardly touches the DRAM pins at all. Reading only the top number tells you that you are memory-limited; reading the levels tells you where, and the fixes for an L1-limited kernel and a DRAM-limited one are different. This is exactly the trap day 30's analytical roofline fell into from the other direction, computing DRAM traffic on paper for a kernel whose re-reads never left cache.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), Nsight Compute 2024.3.2, captured 2026-09-01 with sudo ncu --set full --clock-control base. The Speed Of Light section of day 42's two shipped reports, naive and tiled matmul at 512 cubed:

metric naive tiled
Compute (SM) Throughput 60.78% 72.67%
Memory Throughput 60.78% 72.67%
DRAM Throughput 1.25% 1.95%
Duration 1.18 ms 744.64 us

Both kernels are working the L1 and shared-memory paths hard while the DRAM sits nearly idle, which one glance at the DRAM row proves: this matmul's working set lives in cache at this size. Neither is near 100 on any row, and the stall section of the same reports names the queues they wait in. The tiled kernel's higher percentages with a shorter duration is the healthy pattern: closer to the machine's limit, less time spent.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.