← Glossary
CUDA glossaryMemory
CC 7.5

What is a memory transaction on a GPU?

One request the memory system serves for a warp, covering a fixed span of bytes whether or not the warp wanted all of them.

Two counters describe the same load and people mix them up. A request is one warp-wide load instruction arriving at the memory pipe: 32 lanes, one instruction, one request. A transaction is a sector fetch, 32 bytes wide, and one request turns into as many of them as the 32 addresses need. Nsight Compute reports both, as l1tex__t_requests_pipe_lsu_mem_global_op_ld.sum and l1tex__t_sectors_pipe_lsu_mem_global_op_ld.sum, and the ratio between them is the number that tells you whether a kernel's access pattern is paying for itself. NVIDIA names 4 as the optimal ratio for 32 active threads at a 32-bit access size, which is exactly the contiguous case: 32 lanes times 4 bytes is 128 bytes is four sectors.

The request count is the half that does not move. Day 11 profiled the same copy kernel at six strides and got 2,097,152 requests every time. That is worth sitting with, because it is the whole reason coalescing is a memory problem and not an instruction-count problem. Your access pattern cannot change how many instructions issue. It changes only what each one costs, and the cost lands entirely in the sector column.

The cost has a ceiling, and the ceiling arrives earlier than people expect. Sectors per request doubles with the stride until each lane occupies a sector of its own, which happens at stride 8 for 4-byte elements. From there a warp is already asking for the maximum, 32 sectors, and stride 16 and stride 32 cost exactly the same 32. Anyone who explains a strided kernel purely in transactions therefore predicts a flat curve past stride 8. Day 11's timer says otherwise, so a saturated transaction count is a statement about the memory pipe rather than a statement about your runtime.

Two things reduce transactions in practice. Rearranging data so consecutive lanes hold consecutive addresses is the big one, and it is the only one that changes the ratio from 32 back to 4. A vectorized load is the other: reading float4 gives each lane 16 bytes, so two lanes fill a sector and a warp covers the same bytes in a quarter of the requests. It needs a 16-byte aligned pointer, and it reports misaligned address when it does not get one.

Measured

On a Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75, day 11 ran one copy kernel per stride under Nsight Compute. Every stride moved 536,870,912 bytes.

Stride Load requests Load sectors Sectors per request GB/s
1 2,097,152 8,388,608 4 232.9
2 2,097,152 16,777,216 8 87.2
4 2,097,152 33,554,432 16 41.8
8 2,097,152 67,108,864 32 20.1
16 2,097,152 67,108,864 32 14.2
32 2,097,152 67,108,864 32 9.6

Bandwidth comes from the timed run, the counters from a profiler pass on the same card, in the same session, on the same kernel. Four at stride 1 is the model landing on the nose. Thirty-two from stride 8 onward is the model hitting its ceiling, and the three rows that share that ceiling while bandwidth halves twice are the reason this entry and memory coalescing are two different pages.

Reading these counters needs performance-counter access. Plain ncu returned ERR_NVGPUCTRPERM; sudo ncu worked. The cause is the stock driver default RmProfilingAdminOnly: 1 in /proc/driver/nvidia/params, which a hosted tier such as Colab does not let you change.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. This entry stays a draft until a named author and a different named reviewer sign it.