What is a memory transaction on a GPU?
One request the memory system serves for a warp, covering a fixed span of bytes whether or not the warp wanted all of them.
Two counters describe the same load and people mix them up. A request is one warp-wide load instruction arriving at the memory pipe: 32 lanes, one instruction, one request. A transaction is a sector fetch, 32 bytes wide, and one request turns into as many of them as the 32 addresses need. Nsight Compute reports both, as l1tex__t_requests_pipe_lsu_mem_global_op_ld.sum and l1tex__t_sectors_pipe_lsu_mem_global_op_ld.sum, and the ratio between them is the number that tells you whether a kernel's access pattern is paying for itself. NVIDIA names 4 as the optimal ratio for 32 active threads at a 32-bit access size, which is exactly the contiguous case: 32 lanes times 4 bytes is 128 bytes is four sectors.
The request count is the half that does not move. Day 11 profiled the same copy kernel at six strides and got 2,097,152 requests every time. That is worth sitting with, because it is the whole reason coalescing is a memory problem and not an instruction-count problem. Your access pattern cannot change how many instructions issue. It changes only what each one costs, and the cost lands entirely in the sector column.
The cost has a ceiling, and the ceiling arrives earlier than people expect. Sectors per request doubles with the stride until each lane occupies a sector of its own, which happens at stride 8 for 4-byte elements. From there a warp is already asking for the maximum, 32 sectors, and stride 16 and stride 32 cost exactly the same 32. Anyone who explains a strided kernel purely in transactions therefore predicts a flat curve past stride 8. Day 11's timer says otherwise, so a saturated transaction count is a statement about the memory pipe rather than a statement about your runtime.
Two things reduce transactions in practice. Rearranging data so consecutive lanes hold consecutive addresses is the big one, and it is the only one that changes the ratio from 32 back to 4. A vectorized load is the other: reading float4 gives each lane 16 bytes, so two lanes fill a sector and a warp covers the same bytes in a quarter of the requests. It needs a 16-byte aligned pointer, and it reports misaligned address when it does not get one.
Measured
On a Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75, day 11 ran one copy kernel per stride under Nsight Compute. Every stride moved 536,870,912 bytes.
| Stride | Load requests | Load sectors | Sectors per request | GB/s |
|---|---|---|---|---|
| 1 | 2,097,152 | 8,388,608 | 4 | 232.9 |
| 2 | 2,097,152 | 16,777,216 | 8 | 87.2 |
| 4 | 2,097,152 | 33,554,432 | 16 | 41.8 |
| 8 | 2,097,152 | 67,108,864 | 32 | 20.1 |
| 16 | 2,097,152 | 67,108,864 | 32 | 14.2 |
| 32 | 2,097,152 | 67,108,864 | 32 | 9.6 |
Bandwidth comes from the timed run, the counters from a profiler pass on the same card, in the same session, on the same kernel. Four at stride 1 is the model landing on the nose. Thirty-two from stride 8 onward is the model hitting its ceiling, and the three rows that share that ceiling while bandwidth halves twice are the reason this entry and memory coalescing are two different pages.
Reading these counters needs performance-counter access. Plain ncu returned ERR_NVGPUCTRPERM; sudo ncu worked. The cause is the stock driver default RmProfilingAdminOnly: 1 in /proc/driver/nvidia/params, which a hosted tier such as Colab does not let you change.
Related terms
Where you meet this
- Day 11, CUDA memory coalescing, measured, the lesson that owns this term and produced the counters above.
- Day 12, matrix transpose, where one side of every access is strided by construction.
misaligned address, the failure you get when a vectorized load meets an unaligned pointer.
Sources
- Nsight Compute Profiling Guide, on requests, sectors and the optimal ratio: https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html (checked 2026-08-29)
- CUDA Programming Guide, coalesced global memory access: https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/writing-cuda-kernels.html (checked 2026-08-29)
- NVIDIA on
ERR_NVGPUCTRPERMand performance-counter permissions: https://developer.nvidia.com/nvidia-development-tools-solutions-err_nvgpuctrperm-permission-issue-performance-counters (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. This entry stays a draft until a named author and a different named reviewer sign it.