CUDA ERROR
fatal error: cuda_runtime.h: No such file or directory
The compiler has no CUDA include path. On CUDA 12.6, g++ fails at the include while nvcc compiles the same source. Diagnose which compiler ran.
The compiler processing your file has no CUDA include directory on its search path, because it is not nvcc and nobody told it where the toolkit lives.
| Kind | Compile-time, include path. No cudaError code |
| Captured with | CUDA 12.6 (V12.6.85), host g++ 13, 2026-09-01 |
| Source | code/errors/_toolchain/evidence/run-2026-09-01.txt |
Strings a reader might paste
Reproduced by renaming a working CUDA source to .cpp and handing it to g++ with no -I. From code/errors/_toolchain/evidence/run-2026-09-01.txt:
$ g++ -c hello.cu -o hello.o (renamed to hello.cpp first; g++, no -I)
hello.cpp:2:10: fatal error: cuda_runtime.h: No such file or directory
2 | #include <cuda_runtime.h>
| ^~~~~~~~~~~~~~~~
compilation terminated.
exit: 1
The same file, unrenamed and given to nvcc, compiles without an include flag. That is the whole diagnosis in one comparison: nvcc adds the toolkit's include directory automatically and no other compiler does.
Other spellings of the same failure: cc1plus: fatal error: cuda_runtime.h: No such file or directory from the g++ front end, and, from CMake, Unable to find cuda_runtime.h in "<dir>" for _CUDA_INCLUDE_DIRS.
Cause 1: PyTorch load_inline compiled the C++ half with the host compiler
The highest-traffic version of this error, and the one that stops people following GPU MODE lectures. torch.utils.cpp_extension.load_inline compiles your cpp_sources with the host compiler, which has no CUDA include path, so a #include <cuda_runtime.h> in that string fails while the same include in cuda_sources is fine.
load_inline(name="m", cpp_sources=cpp_src, cuda_sources=cuda_src, functions=["f"])
Fix: keep #include <cuda_runtime.h> in the cuda_sources string only. If the C++ half genuinely needs it, pass extra_include_paths=[f"{os.environ['CUDA_HOME']}/include"], having set CUDA_HOME first.
Cause 2: compiling a CUDA-including file with g++ instead of nvcc
Exactly what our transcript does, on purpose. It also happens by accident whenever a build system decides a file is plain C++, which is what a wrong file extension or a missing language declaration will do.
Fix: compile it with nvcc, or tell g++ where to look and what to link:
g++ -I/usr/local/cuda/include hello.cpp -L/usr/local/cuda/lib64 -lcudart -o hello
Both halves matter. The include path fixes the compile, and -lcudart fixes the link error you meet immediately afterwards.
Cause 3: the toolkit is not installed
The driver-only machine, wearing a different message. If nvcc is also missing, this is the same problem as the nvcc: command not found page and that is the one to read.
Fix: install the toolkit.
Confirm which one you have
find /usr/local/cuda* -name cuda_runtime.h 2>/dev/null | head
If that prints a path, the toolkit is installed, the problem is the search path, and you have just found the value to pass to -I. If it prints nothing, go and install the toolkit; there is nothing to point at yet.
Prevention
- Let one tool own CUDA compilation. Mixed builds where some translation units go to nvcc and some to the host compiler are where this error lives, and CMake's
enable_language(CUDA)exists so you do not have to hand-manage it. - Set
CUDA_HOMEonce, in your profile, and derive include and library paths from it rather than repeating a literal path in five build files.
Related errors
nvcc: command not found, but nvidia-smi works: the same missing toolkit, discovered one step earlier.No CMAKE_CUDA_COMPILER could be found: the build-system version of not knowing where CUDA is.a value of type "void *" cannot be assigned: the next thing that stops the build once the header is found.
The lesson
Day 3 owns installation and the include paths; start at /setup/install-cuda. Day 67 moves the whole thing into CMake so the paths stop being your problem, and day 87 covers the PyTorch extension route in cause 1. The full code table is on the error hub.
Written 2026-09-01. Compiler output from code/errors/_toolchain/evidence/run-2026-09-01.txt, captured 2026-09-01 with CUDA 12.6 (V12.6.85) and host g++ 13 on ornn-internal-t4-spot-5v08. This is a compile-time error, so no GPU was needed to reproduce it. Author and reviewer: not yet assigned; this page does not publish until both are named.