CUDA ERROR
misaligned address: causes and fix
A float4 load needs 16-byte alignment. On a T4, in+1 returned CUDA error 716 and compute-sanitizer named the address, size and source line.
A load or store of N bytes needs an address that is a multiple of N, and yours was not.
| Enum | cudaErrorMisalignedAddress |
| Code | 716 |
| Source | https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29) |
Strings a reader might paste
- Runtime:
misaligned address - PyTorch:
CUDA error: misaligned addresswith the async suffix; 716 is in PyTorch's async list (https://github.com/pytorch/pytorch/blob/main/c10/cuda/CUDAMiscFunctions.cpp , checked 2026-08-29). - memcheck:
Invalid __global__ read of size 16 bytesfollowed byis misaligned(real capture below).
Cause 1: a float4 load from a pointer that is not 16-byte aligned
The day 44 optimization biting back. cudaMalloc returns well-aligned memory, so the base is fine; the offset is what breaks. Our repro is the two lines everyone writes:
__global__ void loadFloat4(const float* in, float4* out) {
const float4* v = reinterpret_cast<const float4*>(in + 1);
out[0] = v[0]; // in + 1 is 4-byte aligned, float4 wants 16
}
From code/errors/misaligned-address/evidence/run-2026-09-01.txt, Tesla T4 (sm_75), driver 595.84, CUDA 12.6:
cudaMalloc in -> 0 (cudaSuccess: no error)
cudaMalloc out -> 0 (cudaSuccess: no error)
cudaDeviceSynchronize -> 716 (cudaErrorMisalignedAddress: misaligned address)
Fix: make the offset a multiple of four elements, or handle the ragged head and tail with scalar loads. Day 12 and day 11 explain why the vectorized load was tempting in the first place.
Cause 2: a reinterpret_cast inside shared memory
extern __shared__ char smem[]; then casting a byte offset that is not a multiple of the target type's size:
extern __shared__ char smem[];
double* d = reinterpret_cast<double*>(smem + 4); // 4 is not a multiple of 8
Fix: lay the shared memory block out largest type first, or align the offset explicitly. Not captured on hardware yet; the mechanism is the same alignment rule as cause 1.
Cause 3: a packed struct read through a wider type
A struct with a packed attribute or a mixed-size layout read as a whole. Fix: alignas(16) on the struct, or read the fields separately. Also not separately captured; same rule.
Confirm it: the measured sanitizer run
This is the error where compute-sanitizer earns its keep, because it names the address and how far off it is. The real output from our repro, same evidence file, trimmed to the lines that matter:
========= Invalid __global__ read of size 16 bytes
========= at loadFloat4(const float *, float4 *)+0x20
========= by thread (0,0,0) in block (0,0,0)
========= Address 0x7866e3000004 is misaligned
========= and is inside the nearest allocation at 0x7866e3000000 of size 256 bytes
Read it bottom up: the allocation starts at ...000, the access hit ...004, and 4 is not a multiple of 16. The "inside the nearest allocation" line is what separates this from an out-of-bounds bug; the pointer is valid, only its alignment is wrong.
One measured subtlety from the same run: under the sanitizer, the API call did not report 716. The tool's final line was
========= Program hit cudaErrorLaunchFailure (error 719) due to "unspecified launch failure" on CUDA API call to cudaDeviceSynchronize.
So the plain run says 716 and the sanitized run says 719 for the same fault on the same card. If your CI logs disagree with your terminal, this is why; see unspecified launch failure.
In cuda-gdb the signal is CUDA_EXCEPTION_6, Warp Misaligned Address for local and shared segments (https://docs.nvidia.com/cuda/cuda-gdb/index.html , checked 2026-08-29).
Prevention
- Vectorize only when the base and the per-thread offset are both provably multiples of the vector width; state the precondition in a comment next to the cast.
- In shared memory, order the layout largest type first.
Related errors
an illegal memory access was encountered(700): the pointer is bad, not just its alignment.unspecified launch failure(719): what this same fault surfaced as under compute-sanitizer.- 717,
operation not supported on global/shared address space: the address-space cousin, no page yet.
The lesson
Day 44 owns the vectorized-load optimization this error comes out of, and day 11 covers why 128-byte transactions make float4 attractive at all. The full code table is on the error hub; toolchain-side failures start at /setup/install-cuda.
Written 2026-09-01. Transcripts from code/errors/misaligned-address/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.