← Glossary
CUDA glossaryMemory
CC 7.5

What is global memory in CUDA?

The GPU's main DRAM, visible to every thread, and the slowest thing a kernel touches.

Global memory is what cudaMalloc hands you a pointer into. Every thread in every block of every kernel can read and write it, it outlives the launch, and it is the only device memory the host can reach with cudaMemcpy. That reach is the whole reason it exists, and the reason it is slow: it sits off the SM, past the caches, at the bottom of the memory hierarchy below registers and shared memory. The pointer means nothing to your process, which is why reading it on the host is a segmentation fault and not a slow read.

There is no single number for how fast it is, and quoting one is the most common way to be wrong about it. Global memory is served in 32-byte sectors, so what you get depends on how many of the bytes in a sector your warp actually wanted. Day 11 copied the identical 536,870,912 bytes seven different ways on one T4 and the result ranged from 232.9 GB/s down to 9.6. Same card, same instruction count, same bytes. The only variable was the order the addresses arrived in, which is coalescing.

Everything between the SM and the DRAM exists to soften that. On the T4 the L2 is 4096 KiB and every DRAM access goes through it, so a pattern that revisits a cache line pays once instead of twice. That is also why the transaction model stops predicting the curve past stride 8: transactions saturate at one sector per lane, but the lanes keep spreading across more 128-byte lines and then more DRAM pages, and the L2 hit rate falls with them.

The capacity limit is the other half of working with it. A T4 reports 14912 MiB of global memory to cudaGetDeviceProperties against the 15360 MiB nvidia-smi prints for the board, so budget from the smaller number. Allocations that do not fit return cudaErrorMemoryAllocation rather than crashing, and an unchecked one gives you a null pointer that surfaces much later as an illegal memory access was encountered. Check every allocation. Day 5 hands you the macro that does it.

Measured

On a Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75:

What Value Where
Global memory reported to the runtime 14912 MiB day 3
L2 cache 4096 KiB day 3
Copy bandwidth, coalesced 232.9 GB/s day 11
Copy bandwidth, stride 32 9.6 GB/s day 11
Copy bandwidth, 32 contiguous elements per thread 32.6 GB/s day 11

Those bandwidth figures are what the card delivered on a copy kernel, not a spec-sheet peak, and the two are not the same number. Only one of them ran. The 24x spread between the coalesced and stride-32 rows is the entire argument for reading the rest of this cluster: a kernel that is bandwidth-bound, which most simple kernels are, has its speed set here and nowhere else.

Diagram

memory-hierarchy-svg, with global memory as the widest and lowest band. Registers and shared memory sit inside the SM box, L1 beside shared memory, L2 as a single band spanning all SMs, and global memory below it as off-chip DRAM. Each band is labelled with the T4's measured figure where the course has one.

Alt text: "The CUDA memory hierarchy on a Tesla T4. Registers and shared memory are per SM, the 4096 KiB L2 is shared by the whole device, and global memory is 14912 MiB of off-chip DRAM that delivered between 9.6 and 232.9 GB/s depending on the access pattern."

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. This entry stays a draft until a named author and a different named reviewer sign it.