LIBRARY / ERRORS

CUDA errors

23tested fixes and clear checks
ERR

BUILD AND RUN

Find the message you saw

  1. 01
    a PTX JIT compilation failed: causes and fixThe driver understood the PTX version but could not assemble its contents. A bad inline register reproduced CUDA error 218 at launch on a T4.
    reproduced
  2. 02
    a value of type "void *" cannot be assignednvcc compiles .cu files as C++, which will not convert void*. On CUDA 12.6, the diagnostic said cannot be used to initialize, not cannot be assigned.
    reproduced
  3. 03
    an illegal memory access was encountered: CUDA error 700Error 700 is asynchronous and sticky. A T4 transcript shows the launch reporting cudaSuccess and an innocent cudaMemcpy taking the blame.
    draft
  4. 04
    calling a __host__ function from a __device__ functionAn unannotated function is host-only, so a kernel cannot call it. The CUDA 12.6 capture names __global__, not __device__, and raises two errors.
    reproduced
  5. 05
    device-side assert triggered: causes and fixCUDA error 710 names the file, line, block and thread. A T4 capture also shows why a build with -DNDEBUG can pass silently.
    reproduced
  6. 06
    ERR_NVGPUCTRPERM: causes and fixNsight Compute fails for normal users when RmProfilingAdminOnly is 1. On a T4 with ncu 2024.3.2, sudo worked; the module option is the durable fix.
    draft
  7. 07
    fatal error: cuda_runtime.h: No such file or directoryThe compiler has no CUDA include path. On CUDA 12.6, g++ fails at the include while nvcc compiles the same source. Diagnose which compiler ran.
    reproduced
  8. 08
    initialization error: CUDA error 3CUDA error 3 means the runtime could not build a usable context. On a T4, fork() after CUDA initialization failed only in the child.
    reproduced
  9. 09
    invalid argument: causes and fixcudaErrorInvalidValue is code 1. On a T4, a null destination and 53248 bytes of dynamic shared memory failed; the wrong cudaMemcpyKind did not.
    draft
  10. 10
    invalid configuration argument: causes and fixCUDA error 9 means the launch geometry is illegal: a zero grid, a block over 1024 threads, or a grid dimension past 65535.
    draft
  11. 11
    invalid device function: causes and fixCUDA error 98 means the kernel symbol is missing for this device. On a T4, the classic no-rdc reproduction still linked and ran under CUDA 12.6.
    reproduced
  12. 12
    invalid device ordinal: causes and fixCUDA error 101 means the requested device index does not exist. On a one-GPU T4 the error was not sticky; a valid index still worked next.
    reproduced
  13. 13
    invalid device symbol: causes and fixCUDA error 13 means the runtime did not recognise a device symbol. On a T4, the often-blamed ampersand succeeded and a host array failed.
    reproduced
  14. 14
    misaligned address: causes and fixA float4 load needs 16-byte alignment. On a T4, in+1 returned CUDA error 716 and compute-sanitizer named the address, size and source line.
    reproduced
  15. 15
    No CMAKE_CUDA_COMPILER could be found: the CMake CUDA familyFour CMake messages, four causes. On CMake 4.4 we got Failed to find nvcc, and one attempt to scrub the environment did not reproduce it.
    reproduced
  16. 16
    no CUDA-capable device is detected: causes and fixCUDA error 100 often means something hid the GPU rather than broke it. On a T4, one empty environment variable made a working card disappear.
    reproduced
  17. 17
    no kernel image is available for execution on the devicenvcc built for one architecture and the GPU is another. On a T4, an sm_80 binary returned CUDA error 209 at launch and success at the later sync.
    reproduced
  18. 18
    nvcc fatal: Unsupported gpu architectureYour nvcc and -arch flag disagree. On CUDA 12.6, sm_30 failed with nvcc's fatal line while deprecated sm_60 still compiled.
    reproduced
  19. 19
    nvcc: command not found, but nvidia-smi worksThe driver and toolkit are separate installs. On CUDA 12.6, the shell and env printed different messages for the same missing binary; both exited 127.
    reproduced
  20. 20
    out of memory: cudaMalloc returned CUDA error 2CUDA error 2 comes from cudaMalloc, not PyTorch. On a T4, a failed allocation left free memory unchanged and size_t arithmetic did not wrap.
    reproduced
  21. 21
    the provided PTX was compiled with an unsupported toolchain.The driver JIT is older than the toolkit that emitted the PTX. Measured with CUDA 13.3 PTX on a 13.2-capable driver while the matching cubin ran.
    reproduced
  22. 22
    too many resources requested for launch: causes and fixCUDA error 701 means a launch exceeds SM resources. On a T4, two classic triggers returned error 1 instead; 62 registers at 1024 threads was legal.
    reproduced
  23. 23
    unspecified launch failure: causes and fixCUDA error 719 is 700 with less information. On a T4, the classic stack-smash reported 700; the captured 719 came from compute-sanitizer.
    reproduced