What is local memory in CUDA?
Per-thread storage that lives off chip in device DRAM despite its name, holding whatever the compiler cannot keep in registers.
The name is the problem, and NVIDIA says so itself: "Local memory is so named because its scope is local to the thread, not because of its physical location. In fact, local memory is off-chip. Hence, access to local memory is as expensive as access to global memory." Nothing about it is near your thread except the address space. It is the same DRAM global memory sits in, reached through the same caches, at the same cost per access. The only thing local is who can see it, which is one thread and nobody else.
Values land there for two reasons and the best practices guide names both: "large structures or arrays that would consume too much register space and arrays that the compiler determines may be indexed dynamically". The second one catches people who write correct C. On a CPU float v[64] is memory, &v[k] is a pointer and k can be anything, because the hardware does the address arithmetic at run time. A register has no address and no instruction reads register number k for a runtime k, so an array survives in registers only while the compiler can rewrite every access into a named register. Give it a literal index after unrolling and the array stops being an array. Give it a variable and the array becomes DRAM.
The third route in is spilling, and it is worth keeping separate. A spill means the compiler wanted registers and could not have them; a dynamically indexed array means it never tried. Both report as local memory bytes and only the -Xptxas -v report splits them, which matters because more registers fixes one and does nothing for the other.
Measured
On a Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75, day 17 ran two kernels that compute the same 64-tap filter over one input at 256 threads a block. One indexes its tap array with the loop counter, the other adds a shift that arrives as a kernel argument.
| Kernel | Registers | Local memory | Occupancy | Time |
|---|---|---|---|---|
decayWide |
72 | 0 B | 75.0 % | 0.164 ms |
decayWideIndexed |
69 | 256 B | 75.0 % | 1.556 ms |
shift is passed as 0, so the two kernels do the same arithmetic in the same order, and the run confirmed it by reporting no differing elements between them. Occupancy is identical. Register counts are within three of each other. The only change is that 64 floats, 256 bytes of them, moved off the chip, and that cost 9.5x. This is the number the name hides, and it is why "local" deserves to be read as "in DRAM, one thread's copy".
Code
From code/day17-registers/registers.cu. One expression is the whole difference between the two rows above.
float acc = 0.0f;
#pragma unroll
for (int k = kWideTaps - 1; k >= 0; --k) {
acc = acc * kDecay + window[(k + shift) & (kWideTaps - 1)];
}
Drop + shift and every index is a literal once the loop unrolls, so window stays in registers. Keep it and the compiler cannot name the element being read, so window needs an address, and an address means memory.
Diagram
memory-hierarchy-svg, with local memory drawn inside the same off-chip DRAM band as global memory rather than beside the register file, and a dashed per-thread boundary around one slice of it.
Alt text: "Local memory sits in the same off-chip DRAM as global memory, not on the SM. What makes it local is that one thread can see its slice, not that it is close to that thread."
Related terms
Where you meet this
- Day 17, registers, local memory and spills, the lesson that owns this term and measured the 9.5x above.
- Day 13, shared memory and tiling, the place to move an array that will not fit in registers and is read by more than one thread.
- Day 11, memory coalescing, measured, for what DRAM actually charges once you are paying DRAM prices.
too many resources requested for launch, the neighbouring failure when the compiler cannot fit the kernel at all.
Sources
- CUDA C++ Best Practices Guide, 10.2.4 "Local Memory", for the off-chip statement and the two cases that send an array there: https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html (checked 2026-08-30)
- NVIDIA CUDA Compiler Driver NVCC, 8.4 "Printing Code Generation Statistics", for the stack frame and spill fields in the report: https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html (checked 2026-08-30)
- CUDA Programming Guide, 5.4.3.2 "Launch Bounds", which warns that a tight bound increases local memory usage: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/cpp-language-extensions.html (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. This entry stays a draft until a named author and a different named reviewer sign it.