CUDA ERROR
unspecified launch failure: causes and fix
CUDA error 719 is 700 with less information. On a T4, the classic stack-smash reported 700; the captured 719 came from compute-sanitizer.
The kernel died and the driver could not attribute the fault to an instruction, so the context is now unusable and every later call in the process fails too.
| Enum | cudaErrorLaunchFailure |
| Code | 719 |
| Source | https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29) |
Strings a reader might paste
- Runtime:
unspecified launch failure - PyTorch:
CUDA error: unspecified launch failurewith the "might be asynchronously reported" suffix; 719 is one of five codes PyTorch tags that way, alongside 700, 710, 715 and 716 (https://github.com/pytorch/pytorch/blob/main/c10/cuda/CUDAMiscFunctions.cpp , functionget_cuda_async_error_suffix, checked 2026-08-29).
The measured reality first: you may not get 719 when the folklore says you will
The textbook trigger for this error is a hardware stack overflow. We built it: a __device__ function with a float pad[4096] local array recursing a million deep. On this card it does not report 719. From code/errors/unspecified-launch-failure/evidence/run-2026-09-01.txt, Tesla T4 (sm_75), driver 595.84, CUDA 12.6:
cudaMalloc -> 0 (cudaSuccess: no error)
peek after launch -> 0 (cudaSuccess: no error)
cudaDeviceSynchronize -> 700 (cudaErrorIllegalAddress: an illegal memory access was encountered)
cudaFree after the failure -> 700 (cudaErrorIllegalAddress: an illegal memory access was encountered)
The one place a 719 did surface in the same session was under the sanitizer. Running compute-sanitizer on the misaligned-load repro, the tool reported the fault precisely and then the API call surfaced as 719 rather than the 716 the plain run returns (code/errors/misaligned-address/evidence/run-2026-09-01.txt):
========= Program hit cudaErrorLaunchFailure (error 719) due to "unspecified launch failure" on CUDA API call to cudaDeviceSynchronize.
Which code you get for one physical fault depends on the card, the driver and whether a tool is attached. So treat 719 as an illegal memory access was encountered with less information, and diagnose it the same way.
Cause 1: the same three bugs as error 700, reported imprecisely
An out-of-bounds index, a host pointer on the device, a use after free. Older cards, virtualised setups, WSL2 and, as measured above, an attached sanitizer can all fold a precise fault into this one code. Start from the 700 page's ranked causes; they are this page's causes.
Cause 2: stack overflow inside the kernel
Deep recursion or a large local array per thread. Our repro is exactly this shape:
__device__ float blow(int depth) {
float pad[4096];
pad[threadIdx.x] = static_cast<float>(depth);
if (depth > 0) { pad[0] += blow(depth - 1); }
return pad[threadIdx.x] + pad[0];
}
On this T4 it surfaced as 700 (transcript above); on other stacks it is the classic 719. Fix either way: bound the recursion, move the array to shared or global memory, or raise cudaLimitStackSize. In cuda-gdb this cause identifies itself as CUDA_EXCEPTION_9, Warp Hardware Stack Overflow (https://docs.nvidia.com/cuda/cuda-gdb/index.html , checked 2026-08-29), which is what separates it from cause 1.
Cause 3: the display driver reset the GPU
On Windows the watchdog kills any kernel that holds the display GPU past its timeout. That case has its own code, 702, and its own page once captured; on a headless card like this T4 it cannot happen.
Confirm which one you have
compute-sanitizer --tool memcheck ./a.out
If memcheck names an invalid access, it is cause 1 and the 700 page's fixes apply; note from the transcript above that the sanitizer may relabel the API-level code to 719 while doing so. If memcheck is clean and cuda-gdb reports CUDA_EXCEPTION_9, it is cause 2.
One more thing the day 6 measurements make concrete: 719 is sticky. After it, every call in the process returns an error until the process restarts, so do not chase the second, third and fourth messages.
Prevention
- The
CUDA_CHECKmacro pluscudaGetLastError()after every launch, per day 6 and the error checking discipline. - Keep per-thread local arrays small; a kilobyte-scale array per thread is a design smell before it is a crash.
Related errors
an illegal memory access was encountered(700): the precise form; on this card, also what the imprecise folklore cases actually returned.misaligned address(716): a specific fault that the sanitizer surfaced as 719 in our capture.device-side assert triggered(710): a deliberate death, with a message that names the thread.- Siblings 714 (hardware stack error), 715 (illegal instruction), 717, 718: same class of bug, rarer labels, no separate pages.
The lesson
Day 6 measures when errors arrive and what sticky means; day 61 finds the underlying bugs with compute-sanitizer. The full code table is on the error hub, and a machine that fails before any launch belongs at /setup/install-cuda.
Written 2026-09-01. Transcripts from code/errors/unspecified-launch-failure/evidence/run-2026-09-01.txt and code/errors/misaligned-address/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.