CUDA errorsdraft
Documented

CUDA ERROR

an illegal memory access was encountered: CUDA error 700

Error 700 is asynchronous and sticky. A T4 transcript shows the launch reporting cudaSuccess and an innocent cudaMemcpy taking the blame.

A thread in your kernel read or wrote an address outside any allocation it is allowed to touch, and the whole CUDA context died with it.

Enum cudaErrorIllegalAddress
Code 700
Source https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29)

Strings a reader might paste

The error arrives late, and it names the wrong line

This is the fact that makes 700 confusing, so here is a measured run rather than a claim. The transcript below is from code/day06-error-checking/evidence/run-2026-08-30.txt, captured on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6. The kernel writePastEnd writes one float 2^28 elements past its buffer. The cudaMemcpy that follows it is entirely correct.

part 3: an error that arrives late
  peek, straight after the launch  cudaSuccess                    no error
  the next cudaMemcpy returned     cudaErrorIllegalAddress        an illegal memory access was encountered

The launch itself reported nothing, because the launch only queues work. The fault happened on the device, and the host heard about it from the next runtime call that looked. Whatever line your error message names, the bug is upstream of it.

The same transcript shows the second property, stickiness. After the fault, every call in the process returns 700, including calls that cannot possibly be wrong:

part 4: what clearing a sticky error buys you
  cudaGetLastError returned        cudaErrorIllegalAddress        an illegal memory access was encountered
  peek, right after clearing it    cudaSuccess                    no error
  a fresh cudaMalloc returned      cudaErrorIllegalAddress        an illegal memory access was encountered
  cudaFree(d_in) returned          cudaErrorIllegalAddress        an illegal memory access was encountered
  cudaFree(d_out) returned         cudaErrorIllegalAddress        an illegal memory access was encountered

cudaGetLastError clears the stored status and changes nothing, because the context itself is gone. Read the first message, ignore every message after it, restart the process.

Cause 1: an index off the end of a device array

The most common by a distance. Every "my kernel worked at N=1024 and died at N=1000" report is this.

__global__ void scale(float* a, int n) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;
    a[i] *= 2.0f;                        // no bounds check
}
scale<<<(n + 255) / 256, 256>>>(d_a, n); // n = 1000 launches 1024 threads

Fix: if (i < n). Day 8 makes the bounds check part of the loop shape so it cannot be forgotten. The transcript above was produced by exactly this class of bug, built on purpose.

Cause 2: a host pointer on the device, or the reverse

A malloc pointer passed to a kernel compiles fine and dies at run time. So does dereferencing a cudaMalloc pointer on the host. Fix: one naming convention, h_ and d_, applied everywhere. That this bites real people is documented at https://www.reddit.com/r/CUDA/comments/1eov1et/racking_my_brain_with_an_odd_access_violation/ (checked 2026-08-29).

Cause 3: use after free

A pointer freed in one iteration and launched on in the next. Fix: set the pointer to nullptr after cudaFree and check it before launch.

Confirm which one you have

The reporting line lies, so do not read the stack trace; run the sanitizer:

compute-sanitizer --tool memcheck ./a.out

memcheck names the memory space, the access size, the kernel, the source line and the thread, in the shape NVIDIA documents under "Understanding Memcheck Errors" at https://docs.nvidia.com/compute-sanitizer/ComputeSanitizer/index.html (checked 2026-08-29). For this error the address line reads is out of bounds. We have not pasted a captured memcheck run here because this page's transcripts came from a node where the sanitizer pass has not run yet; the invocation and output shape above are NVIDIA's documented ones.

CUDA_LAUNCH_BLOCKING=1 makes the reporting call the real one, at a large speed cost. NVIDIA's wording: "Disabling asynchronous execution results in slower execution but is useful for debugging" (https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/environment-variables.html , checked 2026-08-30). Never benchmark with it set.

Prevention

  • Wrap every call in the CUDA_CHECK macro and follow every launch with cudaGetLastError() plus a checked cudaDeviceSynchronize. Day 6 builds the macro and shows what each of the two lines catches.
  • Bounds-check every global memory write against n, not against the grid.

Related errors

  • unspecified launch failure (719): treat it as 700 with less information; same causes, a card or driver that could not attribute the fault.
  • invalid argument (1): rejected on the host, synchronously, before anything runs. If your error is 1, the failing call really is the one reported.
  • invalid configuration argument (9): the launch itself was refused, so the kernel never ran at all.

The lesson

Day 6 builds this fault on purpose and measures who sees it when; day 61 finds three memory bugs like it with compute-sanitizer. The full code table lives on the error hub. If your machine cannot run CUDA at all, start at /setup/install-cuda.


Written 2026-09-01. The quoted transcript is code/day06-error-checking/evidence/run-2026-08-30.txt, captured 2026-08-30 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.