LIBRARY / ERRORS
CUDA errors
23tested fixes and clear checks
ERR
BUILD AND RUN
Find the message you saw
- 01a PTX JIT compilation failed: causes and fixThe driver understood the PTX version but could not assemble its contents. A bad inline register reproduced CUDA error 218 at launch on a T4.reproduced
- 02a value of type "void *" cannot be assignednvcc compiles .cu files as C++, which will not convert void*. On CUDA 12.6, the diagnostic said cannot be used to initialize, not cannot be assigned.reproduced
- 03an illegal memory access was encountered: CUDA error 700Error 700 is asynchronous and sticky. A T4 transcript shows the launch reporting cudaSuccess and an innocent cudaMemcpy taking the blame.draft
- 04calling a __host__ function from a __device__ functionAn unannotated function is host-only, so a kernel cannot call it. The CUDA 12.6 capture names __global__, not __device__, and raises two errors.reproduced
- 05device-side assert triggered: causes and fixCUDA error 710 names the file, line, block and thread. A T4 capture also shows why a build with -DNDEBUG can pass silently.reproduced
- 06ERR_NVGPUCTRPERM: causes and fixNsight Compute fails for normal users when RmProfilingAdminOnly is 1. On a T4 with ncu 2024.3.2, sudo worked; the module option is the durable fix.draft
- 07fatal error: cuda_runtime.h: No such file or directoryThe compiler has no CUDA include path. On CUDA 12.6, g++ fails at the include while nvcc compiles the same source. Diagnose which compiler ran.reproduced
- 08initialization error: CUDA error 3CUDA error 3 means the runtime could not build a usable context. On a T4, fork() after CUDA initialization failed only in the child.reproduced
- 09invalid argument: causes and fixcudaErrorInvalidValue is code 1. On a T4, a null destination and 53248 bytes of dynamic shared memory failed; the wrong cudaMemcpyKind did not.draft
- 10invalid configuration argument: causes and fixCUDA error 9 means the launch geometry is illegal: a zero grid, a block over 1024 threads, or a grid dimension past 65535.draft
- 11invalid device function: causes and fixCUDA error 98 means the kernel symbol is missing for this device. On a T4, the classic no-rdc reproduction still linked and ran under CUDA 12.6.reproduced
- 12invalid device ordinal: causes and fixCUDA error 101 means the requested device index does not exist. On a one-GPU T4 the error was not sticky; a valid index still worked next.reproduced
- 13invalid device symbol: causes and fixCUDA error 13 means the runtime did not recognise a device symbol. On a T4, the often-blamed ampersand succeeded and a host array failed.reproduced
- 14misaligned address: causes and fixA float4 load needs 16-byte alignment. On a T4, in+1 returned CUDA error 716 and compute-sanitizer named the address, size and source line.reproduced
- 15No CMAKE_CUDA_COMPILER could be found: the CMake CUDA familyFour CMake messages, four causes. On CMake 4.4 we got Failed to find nvcc, and one attempt to scrub the environment did not reproduce it.reproduced
- 16no CUDA-capable device is detected: causes and fixCUDA error 100 often means something hid the GPU rather than broke it. On a T4, one empty environment variable made a working card disappear.reproduced
- 17no kernel image is available for execution on the devicenvcc built for one architecture and the GPU is another. On a T4, an sm_80 binary returned CUDA error 209 at launch and success at the later sync.reproduced
- 18nvcc fatal: Unsupported gpu architectureYour nvcc and -arch flag disagree. On CUDA 12.6, sm_30 failed with nvcc's fatal line while deprecated sm_60 still compiled.reproduced
- 19nvcc: command not found, but nvidia-smi worksThe driver and toolkit are separate installs. On CUDA 12.6, the shell and env printed different messages for the same missing binary; both exited 127.reproduced
- 20out of memory: cudaMalloc returned CUDA error 2CUDA error 2 comes from cudaMalloc, not PyTorch. On a T4, a failed allocation left free memory unchanged and size_t arithmetic did not wrap.reproduced
- 21the provided PTX was compiled with an unsupported toolchain.The driver JIT is older than the toolkit that emitted the PTX. Measured with CUDA 13.3 PTX on a 13.2-capable driver while the matching cubin ran.reproduced
- 22too many resources requested for launch: causes and fixCUDA error 701 means a launch exceeds SM resources. On a T4, two classic triggers returned error 1 instead; 62 registers at 1024 threads was legal.reproduced
- 23unspecified launch failure: causes and fixCUDA error 719 is 700 with less information. On a T4, the classic stack-smash reported 700; the captured 719 came from compute-sanitizer.reproduced