← Glossary
CUDA glossaryMemory
CC 7.5

What are L1 and L2 cache on a GPU?

L1 is per SM and shares its silicon with shared memory; L2 is one cache for the whole device, and every access to DRAM goes through it.

The thing that makes GPU L1 different from the one on your CPU is that you can spend it. On Turing the L1, the texture cache and shared memory are one block of memory on the SM: "Like Pascal and Volta, Turing combines the functionality of the L1 and texture caches into a unified L1/Texture cache." Ask for more shared memory in a kernel and there is less L1 behind it. That also settles a question people ask about the read-only path, because a const __restrict__ pointer compiles to a read-only load and on this card that load is served by L1 rather than by some separate cache sitting next to it.

L2 is the one every SM shares, and it is the last stop before DRAM. Nothing reaches global memory without passing through it, which is why a small read-only table is already cached before anyone moves it into constant memory, and why a pattern that revisits a 128-byte line pays for it once rather than twice. It is also where the transaction model runs out of explaining power: past a certain stride the sector count stops rising while bandwidth keeps falling, and what is falling is locality in L2 and in DRAM rather than the request count.

From compute capability 8.0 you can reserve part of L2 and pin a region in it with cudaAccessPolicyWindow, which is the closest thing to managing L2 by hand. The T4 this course targets is 7.5 and reports nothing to reserve, so on this card the cache you steer yourself is shared memory and nothing else. That is the honest ceiling: for a 7.5 card, treat both caches as things you cooperate with by choosing an access pattern, not as things you configure.

Measured

Before it times anything, day 18 reads the cache configuration off the device. Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75.

What the device reports Value
L2 cache size 4096 KiB
L2 set aside for persisting accesses, maximum 0 B
cudaAccessPolicyWindow size, maximum 0 B

Those two zeros are the reason the L2 residency controls are a forward pointer on this hardware rather than a thing to try. The same run gives a reading on L1: a 32-tap filter, 128 bytes, in __constant__ beat the identical filter in global memory by a ratio of 0.68 at 8,388,608 elements when every lane read the same tap. The gap is that small because the global version was hitting in L1 on every warp after the first, so the dedicated cache had almost nothing left to save.

For the other side of it, day 13 transposed a 2003 by 3001 matrix where each buffer is 22.9 MiB against this card's 4096 KiB of L2. A row-major copy ran at 211.2 GB/s and the naive transpose at 65.2 GB/s, over the same buffers with the same caches in place. A 4096 KiB cache in front of a 22.9 MiB buffer does not rescue a strided access pattern; staging the tile in shared memory took the same transpose to 118.6 GB/s, and that is the lever you have on a 7.5 card.

Diagram

memory-hierarchy-svg, with the unified L1 and shared memory drawn as one block inside each SM box, a single L2 band spanning all 40 SMs and labelled 4096 KiB, and global memory as the DRAM band below it.

Alt text: "On a Tesla T4, L1 and shared memory are one block of on-SM storage, all forty SMs share a single 4096 KiB L2, and every access to off-chip DRAM passes through that L2."

Related terms

Where you meet this

Sources

Byline

Author and reviewer are unassigned. The written and last-checked dates get set when two different named people have signed this entry, and it stays a draft until then.