What is dynamic parallelism in CUDA?
A kernel launching another kernel from the device, which CUDA 12 rebuilt as CDP2.
The feature is what it sounds like: device code calls a kernel with the same <<<>>> syntax the host uses, so a kernel that discovers work can launch more of itself without returning to the CPU. CUDA 12 replaced the original implementation with CDP2, which changed the synchronization rules; the current semantics are in the programming guide (https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/dynamic-parallelism.html , checked 2026-08-29).
This course covers it as a decision rather than a technique, and the decision is: do not reach for it in new code. Each use case that motivated it now has a better-supported answer. Irregular nested work, a kernel finding hot spots and refining them, is served by a worklist the next launch consumes. A pipeline whose stages depend on device-side results is a CUDA graph, and since conditional nodes arrived, even the loop-until-converged case, the last argument for device-side launches, runs as a graph with the decision made on the device. A persistent kernel with grid-wide sync covers the rest. What CDP costs, meanwhile, is real: device-side launch overhead, a separate device runtime, harder profiling, and pending-launch limits that surface as runtime errors under load.
Day 57 measures the strongest CDP use case being served without it, which is the honest way to close the question: if the alternative is measurably fine where CDP was supposed to be necessary, the paragraph above is a conclusion rather than an opinion.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75 -lineinfo, captured 2026-09-01. Day 57 runs a Jacobi solver whose iteration count depends on a device-computed residual, the classic "the GPU must decide" workload, three ways:
| strategy | per iteration (ms) | iterations | residual |
|---|---|---|---|
| host loop, sync each step | 0.549 | 26 | 6.914e-06 |
| graph per iteration | 0.541 | 26 | 6.914e-06 |
| conditional graph, device decides | 0.534 | 26 | 6.914e-06 |
The conditional graph makes the convergence decision on the device, no host round trip and no device-side kernel launch, and it is the fastest of the three with bit-identical convergence. The data-dependent control flow that dynamic parallelism was built for runs here as a graph, on a T4, with conditional-node support confirmed by the same program (CUDART 12060). That is the measured basis for this entry's advice.
Related terms
Where you meet this
- Day 57, graph updates and conditional nodes, the lesson that owns this term's decision and measured the table above.
- Day 56, CUDA graphs, the machinery that replaced most CDP use cases.
- Day 28, cooperative groups, the persistent-kernel alternative.
Sources
- CUDA Programming Guide, "CUDA Dynamic Parallelism", for CDP2's semantics and limits: https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/dynamic-parallelism.html (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.