What is global memory in CUDA?
The GPU's main DRAM, visible to every thread, and the slowest thing a kernel touches.
Global memory is what cudaMalloc hands you a pointer into. Every thread in every block of every kernel can read and write it, it outlives the launch, and it is the only device memory the host can reach with cudaMemcpy. That reach is the whole reason it exists, and the reason it is slow: it sits off the SM, past the caches, at the bottom of the memory hierarchy below registers and shared memory. The pointer means nothing to your process, which is why reading it on the host is a segmentation fault and not a slow read.
There is no single number for how fast it is, and quoting one is the most common way to be wrong about it. Global memory is served in 32-byte sectors, so what you get depends on how many of the bytes in a sector your warp actually wanted. Day 11 copied the identical 536,870,912 bytes seven different ways on one T4 and the result ranged from 232.9 GB/s down to 9.6. Same card, same instruction count, same bytes. The only variable was the order the addresses arrived in, which is coalescing.
Everything between the SM and the DRAM exists to soften that. On the T4 the L2 is 4096 KiB and every DRAM access goes through it, so a pattern that revisits a cache line pays once instead of twice. That is also why the transaction model stops predicting the curve past stride 8: transactions saturate at one sector per lane, but the lanes keep spreading across more 128-byte lines and then more DRAM pages, and the L2 hit rate falls with them.
The capacity limit is the other half of working with it. A T4 reports 14912 MiB of global memory to cudaGetDeviceProperties against the 15360 MiB nvidia-smi prints for the board, so budget from the smaller number. Allocations that do not fit return cudaErrorMemoryAllocation rather than crashing, and an unchecked one gives you a null pointer that surfaces much later as an illegal memory access was encountered. Check every allocation. Day 5 hands you the macro that does it.
Measured
On a Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75:
| What | Value | Where |
|---|---|---|
| Global memory reported to the runtime | 14912 MiB | day 3 |
| L2 cache | 4096 KiB | day 3 |
| Copy bandwidth, coalesced | 232.9 GB/s | day 11 |
| Copy bandwidth, stride 32 | 9.6 GB/s | day 11 |
| Copy bandwidth, 32 contiguous elements per thread | 32.6 GB/s | day 11 |
Those bandwidth figures are what the card delivered on a copy kernel, not a spec-sheet peak, and the two are not the same number. Only one of them ran. The 24x spread between the coalesced and stride-32 rows is the entire argument for reading the rest of this cluster: a kernel that is bandwidth-bound, which most simple kernels are, has its speed set here and nowhere else.
Diagram
memory-hierarchy-svg, with global memory as the widest and lowest band. Registers and shared memory sit inside the SM box, L1 beside shared memory, L2 as a single band spanning all SMs, and global memory below it as off-chip DRAM. Each band is labelled with the T4's measured figure where the course has one.
Alt text: "The CUDA memory hierarchy on a Tesla T4. Registers and shared memory are per SM, the 4096 KiB L2 is shared by the whole device, and global memory is 14912 MiB of off-chip DRAM that delivered between 9.6 and 232.9 GB/s depending on the access pattern."
Related terms
- memory coalescing
- memory transaction
- sector and cache line
- L1 and L2 cache
- GPU RAM
- memory bandwidth
cudaMemcpy
Where you meet this
- Day 5, CUDA vector addition, the first program that allocates it, fills it and reads it back.
- Day 11, CUDA memory coalescing, measured, the lesson that owns this term and measured the bandwidth range above.
- How to set up CUDA, which prints your own card's global memory and L2 size.
out of memory, whencudaMalloccannot find the space.an illegal memory access was encountered, the usual end of an unchecked allocation or an out-of-bounds index.
Sources
- CUDA Programming Guide, on global memory access and coalescing: https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/writing-cuda-kernels.html (checked 2026-08-29)
- CUDA Runtime API, memory management: https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__MEMORY.html (checked 2026-08-29)
- Nsight Compute Profiling Guide, on the 32-byte sector: https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. This entry stays a draft until a named author and a different named reviewer sign it.