What is stream-ordered allocation in CUDA?
Allocating and freeing device memory inside a stream, from a pool, so the call does not synchronize the whole device.
cudaMalloc and cudaFree live outside the stream model: they take effect immediately, which forces the driver to synchronize with in-flight work, and a free must wait until nothing can still touch the pointer. Put them inside a loop and the loop inherits a device-wide stall per iteration. cudaMallocAsync and cudaFreeAsync move the operations into a stream, where they become ordered work items like any launch: the allocation is valid for work after it in the stream, the free happens once prior work completes, and no global synchronization is needed (https://developer.nvidia.com/blog/using-cuda-stream-ordered-memory-allocator-part-1/ , checked 2026-09-01).
The speed comes from a pool. Freed-async memory returns to a cudaMemPool_t rather than to the driver, so the next allocation is a cheap reuse instead of a real allocation. The knob that surprises people is the release threshold: by default the pool gives memory back to the OS at every synchronization, so a loop that syncs each iteration rebuilds the pool each time. Day 55 shows both settings: after a sync the pool held 0 bytes at threshold 0 and 33,554,432 bytes at threshold max, and holding the memory is what makes the fast path fast across sync points.
The honest comparison in day 55's table is not async against sync, it is async against not allocating at all. Hoisting the allocation out of the loop is still the best answer when the size is fixed; cudaMallocAsync earns its place when sizes change per iteration or the code cannot hoist, and it gets within 1 percent of hoisted on this card.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75 -lineinfo, captured 2026-09-01. Day 55's scratch-buffer loop, 100 iterations, all three variants producing identical results:
| loop | us per iteration | vs hoisted |
|---|---|---|
| cudaMalloc and cudaFree per iteration | 856.4 | 2.56x |
| allocation hoisted out of the loop | 334.1 | 1.00x |
| cudaMallocAsync and cudaFreeAsync per iteration | 336.9 | 1.01x |
The nsys API summary puts the cost where the prose claims it: 104 cudaMalloc calls cost 81.143 ms and 104 cudaFree calls cost 49.310 ms in the synchronous loop, which is most of the 2.56x. The stream-ordered loop pays 1 percent over never allocating, while keeping the per-iteration allocation semantics.
Related terms
Where you meet this
- Day 55, stream-ordered allocation, the lesson that owns this term and measured the loop.
- Day 51, CUDA streams, the ordering model the allocator joins.
- Day 56, CUDA graphs, where allocations inside captured work follow the same stream-ordered rules.
Sources
- NVIDIA developer blog, "Using the NVIDIA CUDA Stream-Ordered Memory Allocator", for the semantics and the release threshold: https://developer.nvidia.com/blog/using-cuda-stream-ordered-memory-allocator-part-1/ (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.