← Glossary
CUDA glossaryMemory
CC 7.5

What is pinned memory in CUDA?

Host memory the OS cannot page out, which lets the GPU copy from it by DMA and lets copies overlap kernels.

The GPU's copy engine reads host memory by physical address, and it can only do that safely if the page is guaranteed to stay put. Ordinary malloc memory carries no such guarantee, so a copy from it goes through a hidden staging step: the driver moves your data into a pinned buffer of its own, then DMAs from there. cudaMallocHost (or cudaHostAlloc) gives you memory that is pinned from the start, cutting the extra pass. The 55,040-view Stack Overflow answer covers the mechanism (https://stackoverflow.com/questions/5736968/why-is-cuda-pinned-memory-so-fast , checked 2026-08-29); what it does not carry is a measured table, which is what day 53 adds.

The second effect is quieter and worth more: cudaMemcpyAsync from pageable memory is not actually asynchronous. The staging step runs on the calling thread, so the "async" call blocks for essentially the whole transfer, and every overlap scheme built on it silently serializes. From pinned memory the call really does return immediately, which is what makes copy-compute overlap in streams and the double-buffering of day 54 possible at all.

Two costs, stated. Pinning gigabytes bites the OS: page-locked memory is removed from what the kernel can manage, so allocate pinned buffers for transfer staging, not as a default. And mapped ("zero-copy") memory, pinned memory the kernel dereferences directly over the bus, turns every access into PCIe traffic: day 53 measured a kernel reading its input that way at 9.93x the cost of the device-resident version.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75 -lineinfo, captured 2026-09-01. Day 53's host-to-device cudaMemcpy, mean of 10 runs per cell:

size pageable pinned ratio
1 MiB 3.4 GB/s 9.3 GB/s 2.74x
16 MiB 2.8 GB/s 11.2 GB/s 4.06x
256 MiB 4.4 GB/s 12.2 GB/s 2.81x

The async check at 64 MiB: from pageable memory the cudaMemcpyAsync call itself took 15.295 ms of a 15.387 ms completion, 99.4 percent, blocking in all but name; from pinned memory the call took 0.006 ms of a 5.625 ms completion, 0.1 percent, genuinely asynchronous. Across the ten timed calls the API-side cost was 146.556 ms pageable against 0.085 ms pinned in the nsys trace, a 1717.6x difference in what the calling thread paid.

Diagram

timeline-host-device, preset staged-vs-dma: a pageable copy shown as two bars (staging on the host lane, DMA on the copy lane) against a pinned copy's single DMA bar with the host lane free.

Alt text: "A pageable copy stages through a driver buffer and blocks the host; a pinned copy is one DMA transfer and the host thread is free during it."

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.