← Glossary
CUDA glossaryMemory
CC 7.5

What is a memory sector on a GPU?

A sector is the 32-byte unit the memory system actually fetches, and a 128-byte cache line is four of them.

Nsight Compute defines a sector as an "aligned 32 byte-chunk of memory in a cache line", and it is the granularity you are billed at. Ask for one float and you get 32 bytes. Ask for four floats that share a sector and you still get 32 bytes, which is why coalescing works at all. The same guide states the other half of the arithmetic: in the ideal 32-bit case "every 8 consecutive threads access the same sector", so a warp of 32 lanes covers four sectors and NVIDIA calls four the optimal sectors-per-request ratio.

That gives you a formula nobody writes down. A 32-byte sector holds eight 4-byte floats, so at stride s a warp uses 1 / min(s, elementsPerSector) of every byte fetched, which is 1 / min(s, 8) for floats. Stride 2 wastes half. Stride 4 wastes three quarters. Stride 8 wastes seven eighths, and there it stops getting worse, because at stride 8 each lane already sits in a sector of its own and a 32-lane warp cannot demand more than 32 sectors from one load. The waste saturates at stride 8, not at stride 32, which is the part most explanations get backwards.

Sectors and lines then answer different questions. Sectors set what a transaction costs and they stop moving at stride 8. Lines set what the caches can do for you, and they keep moving: at stride 8 a warp's 32 sectors come from eight 128-byte lines, at stride 32 they come from 32 separate lines, and past that from more DRAM pages. Day 11 measured the consequence directly. Sectors identical, bandwidth halved again and again. If you only count sectors you will predict that the curve flattens, and it does not.

The practical use is sizing. Anything smaller than 32 bytes per warp-step is buying bytes it will throw away, so a struct of arrays beats an array of structs, a padded row that keeps each row 32-byte aligned beats a tight one, and a vectorized load of float4 puts two lanes in each sector instead of eight, which lowers the request count without changing the sector count. Global memory has no mode where you fetch less than a sector.

Measured

On a Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75, day 11 profiled one copy kernel per stride with Nsight Compute. Every variant moved 536,870,912 bytes.

Stride Load requests Load sectors Sectors per request Bytes used per sector GB/s
1 2,097,152 8,388,608 4 32 of 32 232.9
2 2,097,152 16,777,216 8 16 of 32 87.2
4 2,097,152 33,554,432 16 8 of 32 41.8
8 2,097,152 67,108,864 32 4 of 32 20.1
16 2,097,152 67,108,864 32 4 of 32 14.2
32 2,097,152 67,108,864 32 4 of 32 9.6

The first row is the model being exactly right: 32 lanes times 4 bytes is 128 bytes is four sectors, and the profiler says four. The last three rows are the model running out. Same sectors, same bytes used per sector, and bandwidth still falls from 20.1 to 9.6 GB/s, because the lanes are spreading across more 128-byte lines and more DRAM pages. The counters came from sudo ncu; as a normal user the run fails with ERR_NVGPUCTRPERM, which is the stock driver default and the reason a Colab session cannot reproduce this table.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. This entry stays a draft until a named author and a different named reviewer sign it.