← Glossary
CUDA glossaryExecution model
CC 7.5

What is a CUDA graph?

A captured set of kernels and copies with their dependencies, launched as one object so you pay the launch cost once.

The cheapest way to build one is stream capture: bracket your existing launches with cudaStreamBeginCapture and cudaStreamEndCapture, and instead of executing, the work is recorded as a graph, kernels as nodes, the stream and event ordering as edges (https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/cuda-graphs.html , checked 2026-09-01). Instantiate the graph once, then cudaGraphLaunch submits the whole thing per iteration: one API call, one submission, however many nodes. What it eliminates is per-kernel launch overhead; what it cannot do is make any kernel faster or move fewer bytes, which is kernel fusion's job. The two get conflated constantly and they attack different costs.

That framing predicts exactly when graphs pay: short kernels, launched often. A chain whose kernels each run tens of microseconds is mostly submission; the same chain at millisecond scale is mostly work, and shaving submission changes nothing visible. Day 56 measures the whole curve on one chain, and the saved microseconds per iteration are nearly constant while the total grows, which is the entire economics of the feature in one table.

Two costs to respect. Instantiation is front-loaded (80.53 us average on the measured node, roughly the price of a few dozen launches), so a graph launched twice is a loss and a graph launched thousands of times is free. And a graph freezes its parameters at instantiation; changing sizes or pointers per iteration needs cudaGraphExecUpdate or conditional nodes, which is day 57's territory.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75 -lineinfo, captured 2026-09-01. Day 56 captured day 48's four-kernel chain (4 launches recorded, 4 nodes) and submitted 20-iteration batches both ways, mean of 10 runs, per whole-chain iteration:

n 4 launches (us) 1 graph (us) saved (us) ratio
1024 19.753 9.739 10.014 2.03
16384 19.260 10.501 8.760 1.83
262144 45.629 41.552 4.078 1.10
16777216 2406.631 2402.352 4.280 1.00

Same kernels, same bytes, both sides. The saved column is bounded by what submission cost in the first place, about 10 us for four launches here, and the ratio column is that constant saving divided by growing work. At the bottom of the sweep the graph doubles throughput; at the top it vanishes into noise. If your chain lives at the top, fuse or optimize the kernels instead.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.