What is a vectorized load in CUDA?
Reading 8 or 16 bytes per lane with float2 or float4, so one warp covers more bytes per instruction.
A plain float load moves 4 bytes per lane, 128 bytes per warp. Load a float4 instead and each lane moves 16 bytes, so the same warp covers 512 bytes with one instruction. The transactions the memory system serves are the same either way for the same bytes; what shrinks is the instruction count, four times fewer load instructions issued, four times fewer address computations, four times less pressure on the queues that feed the memory pipelines. That matters exactly when instruction issue, not DRAM, is the limit, which is the usual state of a well-tiled kernel.
The requirement is alignment: the compiler emits a wide load only for a pointer it can prove is 16-byte aligned, and the built-in vector types carry that alignment (https://docs.nvidia.com/cuda/cuda-programming-guide/03-advanced/advanced-kernel-programming.html , checked 2026-08-30). cudaMalloc allocations satisfy it; an arbitrary offset into one may not, and a misaligned wide access is a runtime error rather than a slow path.
Day 44's ladder isolates the effect. Steps 5 and 6 compute the same micro-tile; step 6 only widens the shared-memory traffic to float4, cutting shared loads per FMA from 0.2500 to 0.0625 and executed shared-load instructions from 41,943,040 to 16,777,216, a 2.5 to 1 cut. One honest cost from the same run: the widened kernel spends 128 registers per thread against 116, and its transposed fill keeps a deliberate two-way shared-store bank conflict, 2,097,152 of them, which padding would remove and which the lesson leaves as the reader's exercise.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75 -lcublas, captured 2026-09-01. Day 44's step 5 (scalar register tile) against step 6 (same tile, float4 loads), mean of 10 timed runs after 3 warm-ups:
| size | step 5 | step 6 | %cuBLAS |
|---|---|---|---|
| 1024 | 1.429 ms, 1503.1 GFLOP/s | 0.854 ms, 2515.6 GFLOP/s | 34.4 to 57.5 |
| 2048 | 6.011 ms, 2858.2 GFLOP/s | 4.219 ms, 4071.7 GFLOP/s | 47.5 to 67.7 |
Same arithmetic, same tile, same occupancy class; the only change is the width of the loads, and it is worth 1.4x at both sizes. The instruction-count cut is the mechanism, confirmed by the profiler's executed shared-load counts above.
Diagram
coalescing-visualizer, preset vector-width: one warp shown twice against the same 512 bytes, first as four 128-byte instructions, then as one instruction of float4 lanes.
Alt text: "The same warp covers 512 bytes with four scalar load instructions or one float4 load instruction; the bytes moved are identical and the instruction count is a quarter."
Related terms
Where you meet this
- Day 44, optimizing matmul, steps 4 to 6, the lesson that owns this term and produced the ladder above.
- Day 11, memory coalescing, where bytes-per-instruction first separates from bytes-per-transaction.
- Day 15, bank conflicts, the fix for the conflict the widened fill keeps on purpose.
Sources
- CUDA Programming Guide, device memory accesses and built-in vector types, for the size and alignment requirements of wide loads: https://docs.nvidia.com/cuda/cuda-programming-guide/03-advanced/advanced-kernel-programming.html (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.