CUDA ERROR
no CUDA-capable device is detected: causes and fix
CUDA error 100 often means something hid the GPU rather than broke it. On a T4, one empty environment variable made a working card disappear.
The CUDA runtime asked the driver for a list of usable GPUs and got an empty list back.
| Enum | cudaErrorNoDevice |
| Code | 100 |
| Source | https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29) |
Strings a reader might paste
- Runtime:
no CUDA-capable device is detected - Neighbours worth telling apart:
CUDA driver is a stub library(34),CUDA-capable device(s) is/are busy or unavailable(46), andsystem not yet initialized(802). All three mean the driver answered; 100 means it answered "none". - In the same shell,
nvidia-smioften prints its own failure about not being able to communicate with the driver. Ifnvidia-smiworks and your program still gets 100, the machine is fine and something in the process environment is hiding the card. That is cause 1.
The same binary, twice, with one variable changed
From code/errors/no-cuda-capable-device/evidence/run-2026-09-01.txt, Tesla T4 (sm_75), driver 595.84, CUDA 12.6. The program is the same in both halves; only the environment differs:
== normal run ==
cudaGetDeviceCount -> 0 (cudaSuccess: no error)
device count: 1
cudaMalloc -> 0 (cudaSuccess: no error)
== CUDA_VISIBLE_DEVICES="" ./repro ==
cudaGetDeviceCount -> 100 (cudaErrorNoDevice: no CUDA-capable device is detected)
device count: -1
cudaMalloc -> 100 (cudaErrorNoDevice: no CUDA-capable device is detected)
Read the third line of the second block carefully. The program set its counter to -1 before the call, and after a failed cudaGetDeviceCount it is still -1. The runtime did not write a zero; it wrote nothing. Code that prints the count without checking the status therefore prints whatever the variable happened to hold, and a program that loops for (int i = 0; i < count; ++i) over an uninitialised count is a bug waiting for a different machine. Check the cudaError_t, never the out-parameter.
The second useful line is the last one. Once the device list is empty, the next call fails with the same 100 rather than something new. There is nothing to recover from and nothing to retry.
Cause 1: CUDA_VISIBLE_DEVICES is empty or invalid
The measured case. An empty string does not mean "no filter", it means "no devices", and the variable is inherited by every child process, so a job scheduler, a wrapper script or a shell you opened an hour ago can set it without you seeing it.
CUDA_VISIBLE_DEVICES= ./a.out # empty means none
Fix: unset the variable, or set it to a valid index. An index that does not exist on the machine hides everything the same way.
Cause 2: a container with no GPU runtime
The image has the CUDA runtime, the host has the driver, and nothing connected them. Without --gpus all and the NVIDIA container toolkit the container genuinely has no device.
Fix: docker run --gpus all ... and install nvidia-container-toolkit on the host.
Cause 3: the kernel module is not loaded
Usually after a kernel upgrade, or with Secure Boot refusing an unsigned module. The GPU is in the machine and the operating system will not talk to it.
Fix: lsmod | grep nvidia to see whether the module is loaded, then dmesg | grep -i nvidia for the reason it is not. Secure Boot needs the module signed or turned off.
Confirm which one you have
In this order, because each step rules out the one below it:
nvidia-smi
echo "[$CUDA_VISIBLE_DEVICES]"
lsmod | grep nvidia
The brackets in the second command are the entire trick. An empty variable prints as an empty line and looks exactly like an unset variable, so [] versus no output at all is the difference between cause 1 and a machine that is genuinely fine. Our transcript is that experiment run twice on one node.
Prevention
- Never read an out-parameter whose call you did not check. This error is the cleanest demonstration in the whole set that the two are independent.
- Print
cudaGetDeviceCountand the device name at start-up in anything long-running, so the log says which card it got rather than assuming.
Related errors
invalid device ordinal(101): the same variable, one step further along, where devices exist but the index does not.initialization error(3): the driver was there and the context still could not be built.out of memory(2): a device that exists but has no room.
The lesson
Day 3 owns the install and the environment; the setup page at /setup/install-cuda is the fastest route from here. Day 9 shows what the runtime does once it has found a card. The full code table is on the error hub.
Written 2026-09-01. Transcript from code/errors/no-cuda-capable-device/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.