← Glossary
CUDA glossaryMemory
CC 7.5

What is stream-ordered allocation in CUDA?

Allocating and freeing device memory inside a stream, from a pool, so the call does not synchronize the whole device.

cudaMalloc and cudaFree live outside the stream model: they take effect immediately, which forces the driver to synchronize with in-flight work, and a free must wait until nothing can still touch the pointer. Put them inside a loop and the loop inherits a device-wide stall per iteration. cudaMallocAsync and cudaFreeAsync move the operations into a stream, where they become ordered work items like any launch: the allocation is valid for work after it in the stream, the free happens once prior work completes, and no global synchronization is needed (https://developer.nvidia.com/blog/using-cuda-stream-ordered-memory-allocator-part-1/ , checked 2026-09-01).

The speed comes from a pool. Freed-async memory returns to a cudaMemPool_t rather than to the driver, so the next allocation is a cheap reuse instead of a real allocation. The knob that surprises people is the release threshold: by default the pool gives memory back to the OS at every synchronization, so a loop that syncs each iteration rebuilds the pool each time. Day 55 shows both settings: after a sync the pool held 0 bytes at threshold 0 and 33,554,432 bytes at threshold max, and holding the memory is what makes the fast path fast across sync points.

The honest comparison in day 55's table is not async against sync, it is async against not allocating at all. Hoisting the allocation out of the loop is still the best answer when the size is fixed; cudaMallocAsync earns its place when sizes change per iteration or the code cannot hoist, and it gets within 1 percent of hoisted on this card.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75 -lineinfo, captured 2026-09-01. Day 55's scratch-buffer loop, 100 iterations, all three variants producing identical results:

loop us per iteration vs hoisted
cudaMalloc and cudaFree per iteration 856.4 2.56x
allocation hoisted out of the loop 334.1 1.00x
cudaMallocAsync and cudaFreeAsync per iteration 336.9 1.01x

The nsys API summary puts the cost where the prose claims it: 104 cudaMalloc calls cost 81.143 ms and 104 cudaFree calls cost 49.310 ms in the synchronous loop, which is most of the 2.56x. The stream-ordered loop pays 1 percent over never allocating, while keeping the per-iteration allocation semantics.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.