What is kernel launch overhead in CUDA?
The fixed cost of getting a kernel started, which dominates when the kernel itself is short.
Every <<<>>> pays a toll before the first thread runs: the host-side runtime call, the command's trip to the GPU, and the hardware distributing blocks to SMs. The toll is microseconds, which is nothing next to a millisecond kernel and everything next to a microsecond one. Divide to know which side you are on: a kernel that does 2 microseconds of work behind 2.668 microseconds of launch is spending most of its wall clock on ceremony, and no amount of tuning inside the kernel touches that.
Three separate costs hide under this one name, and they scale differently. The per-launch toll is the recurring one. On top of it, the first launch of each kernel in a process pays its own module load, because lazy module loading has been the default since CUDA 12.2 on Linux (https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/lazy-loading.html , checked 2026-08-30): day 9 measured vectorAdd at 1.029 ms cold against 0.786 ms warm on this card. And before any of that, the first CUDA call in the process built the context, 255.966 ms in the same run. Benchmarks that skip warm-up launches bill one or both of these to the kernel.
The fixes follow from the arithmetic. Make launches do more work each, which is what kernel fusion does when it collapses a chain into one launch. Or make launches cheaper in bulk, which is what a CUDA graph does by submitting a captured chain as one object; day 56 is the course's lesson for that path.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75, captured 2026-09-01. Day 48 timed 1000 back-to-back launches of a doNothing kernel per run:
| launch shape | per launch |
|---|---|
doNothing, 1 block of 1 thread |
2.668 us |
doNothing, the chain's own 65536-block grid |
71.668 us |
The second row is the part people miss: an empty kernel is not free at scale, because distributing 65536 blocks costs real time even when every block returns immediately, so "launch overhead" is not one number but a floor plus a grid-size term. The same day swept its four-kernel staged chain down in size and watched the per-launch share approach the floor: at 16,777,216 elements each launch carried 606.606 us of real work; at 1024 elements the whole chain's time divided by its four launches was 4.050 us, within sight of the empty kernel's 2.668, at which point the chain was doing almost nothing but launching.
Diagram
timeline-host-device, preset short-kernel-chain: the host lane issuing four launches whose gaps are wider than the four device-lane kernel bars beneath them, then the same work as one fused bar.
Alt text: "Four short kernels spend more timeline on launch gaps than on work; one fused launch pays the toll once."
Related terms
Where you meet this
- Day 48, kernel fusion, the lesson that owns this term and priced the launches above.
- Day 9, why your GPU code looks slower than your CPU, where the cold-launch and context costs are measured.
- Day 41, Nsight Systems, where launch gaps become visible bars on a timeline.
- /setup/check-cuda-version, because whether lazy loading is your default depends on your toolkit.
Sources
- CUDA Programming Guide, "Lazy Loading", for the first-launch module load and its default since 12.2: https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/lazy-loading.html (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.