What is a CUDA stream?
An ordered queue of GPU work, where two streams may overlap and one stream may not.
Everything you submit to the GPU, kernels and copies alike, goes into some stream, and a stream is a promise about order: items in one stream run in submission order, items in different streams have no ordering at all unless you add one with a CUDA event. That second half is the point. Order is what forbids overlap, so putting independent work in separate streams is how you tell the hardware it is allowed to run a copy under a kernel, or a batch of small kernels beside a long one.
Not asking is the default, because code that never names a stream uses the legacy default stream, and its documented behavior is the trap: work in the legacy stream will not begin until all previously issued work in all other streams has completed, and all other streams wait for it in turn (https://docs.nvidia.com/cuda/cuda-runtime-api/api-sync-behavior.html , checked 2026-08-30). One forgotten default-stream launch, a stray cudaMemcpy, a debug kernel, re-serializes everything around it silently. Day 51 measures exactly this: its two-stream version runs at 2.696 ms, and adding a single legacy-stream launch between the submissions drags the same work back to 4.002 ms, within 10 percent of never having used streams at all. cudaStreamNonBlocking at creation is the opt-out.
The other lesson in the numbers is that streams permit overlap; they do not create resources. Two full-size grids in two streams ran at 0.99 of their summed solo times on day 51, because each grid already filled the card's SMs. Overlap pays when the concurrent work uses different engines (copy against compute) or when neither piece alone fills the machine, and the timeline profiler is how you check which case you are in.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75 -lineinfo, captured 2026-09-01. Day 51 runs a long spin kernel plus a sweep of small kernels, mean of 10 runs:
| arrangement | ms |
|---|---|
| everything in the legacy stream | 4.387 |
| split across two non-blocking streams | 2.696 |
| two streams plus one legacy-stream launch | 4.002 |
| two full-size grids in two streams | 5.146 |
The two-stream row is 1.63x faster than the serialized one, and 1.11x the spin kernel running alone, so the sweep is almost fully hidden under it. The nsys trace of one two-stream pass confirms the mechanism from the GPU's side: kernel durations sum to 4.742 ms while the span from earliest start to latest end is 2.773 ms, a 1.71 concurrency ratio measured off the timeline. The last row is the honest limit: 0.99 of the two kernels' summed solo times, nothing gained, because both grids already fill the card.
Diagram
stream-timeline, preset legacy-trap: three lanes (stream A, stream B, legacy) showing the sweep hiding under the spin kernel, then the same submission with one legacy-stream launch forcing a full drain between them.
Alt text: "Two non-blocking streams let small kernels run under a long one; a single launch into the legacy stream forces everything before it to finish first, undoing the overlap."
Related terms
Where you meet this
- Day 51, CUDA streams, the lesson that owns this term and measured the table above.
- Day 52, events and dependencies, for ordering between streams.
- Day 53, pinned memory, the allocation that makes copy-compute overlap possible at all.
- Day 41, Nsight Systems, the tool that shows whether overlap actually happened.
Sources
- CUDA Runtime API, "API synchronization behavior", for the legacy default stream's ordering rules: https://docs.nvidia.com/cuda/cuda-runtime-api/api-sync-behavior.html (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.