CUDA errorsdraft
Reproduced 2026-09-01

CUDA ERROR

misaligned address: causes and fix

A float4 load needs 16-byte alignment. On a T4, in+1 returned CUDA error 716 and compute-sanitizer named the address, size and source line.

A load or store of N bytes needs an address that is a multiple of N, and yours was not.

Enum cudaErrorMisalignedAddress
Code 716
Source https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29)

Strings a reader might paste

Cause 1: a float4 load from a pointer that is not 16-byte aligned

The day 44 optimization biting back. cudaMalloc returns well-aligned memory, so the base is fine; the offset is what breaks. Our repro is the two lines everyone writes:

__global__ void loadFloat4(const float* in, float4* out) {
    const float4* v = reinterpret_cast<const float4*>(in + 1);
    out[0] = v[0];  // in + 1 is 4-byte aligned, float4 wants 16
}

From code/errors/misaligned-address/evidence/run-2026-09-01.txt, Tesla T4 (sm_75), driver 595.84, CUDA 12.6:

cudaMalloc in                      -> 0 (cudaSuccess: no error)
cudaMalloc out                     -> 0 (cudaSuccess: no error)
cudaDeviceSynchronize              -> 716 (cudaErrorMisalignedAddress: misaligned address)

Fix: make the offset a multiple of four elements, or handle the ragged head and tail with scalar loads. Day 12 and day 11 explain why the vectorized load was tempting in the first place.

Cause 2: a reinterpret_cast inside shared memory

extern __shared__ char smem[]; then casting a byte offset that is not a multiple of the target type's size:

extern __shared__ char smem[];
double* d = reinterpret_cast<double*>(smem + 4);   // 4 is not a multiple of 8

Fix: lay the shared memory block out largest type first, or align the offset explicitly. Not captured on hardware yet; the mechanism is the same alignment rule as cause 1.

Cause 3: a packed struct read through a wider type

A struct with a packed attribute or a mixed-size layout read as a whole. Fix: alignas(16) on the struct, or read the fields separately. Also not separately captured; same rule.

Confirm it: the measured sanitizer run

This is the error where compute-sanitizer earns its keep, because it names the address and how far off it is. The real output from our repro, same evidence file, trimmed to the lines that matter:

========= Invalid __global__ read of size 16 bytes
=========     at loadFloat4(const float *, float4 *)+0x20
=========     by thread (0,0,0) in block (0,0,0)
=========     Address 0x7866e3000004 is misaligned
=========     and is inside the nearest allocation at 0x7866e3000000 of size 256 bytes

Read it bottom up: the allocation starts at ...000, the access hit ...004, and 4 is not a multiple of 16. The "inside the nearest allocation" line is what separates this from an out-of-bounds bug; the pointer is valid, only its alignment is wrong.

One measured subtlety from the same run: under the sanitizer, the API call did not report 716. The tool's final line was

========= Program hit cudaErrorLaunchFailure (error 719) due to "unspecified launch failure" on CUDA API call to cudaDeviceSynchronize.

So the plain run says 716 and the sanitized run says 719 for the same fault on the same card. If your CI logs disagree with your terminal, this is why; see unspecified launch failure.

In cuda-gdb the signal is CUDA_EXCEPTION_6, Warp Misaligned Address for local and shared segments (https://docs.nvidia.com/cuda/cuda-gdb/index.html , checked 2026-08-29).

Prevention

  • Vectorize only when the base and the per-thread offset are both provably multiples of the vector width; state the precondition in a comment next to the cast.
  • In shared memory, order the layout largest type first.

Related errors

The lesson

Day 44 owns the vectorized-load optimization this error comes out of, and day 11 covers why 128-byte transactions make float4 attractive at all. The full code table is on the error hub; toolchain-side failures start at /setup/install-cuda.


Written 2026-09-01. Transcripts from code/errors/misaligned-address/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.