CUDA errorsdraft
Reproduced 2026-09-01

CUDA ERROR

invalid device ordinal: causes and fix

CUDA error 101 means the requested device index does not exist. On a one-GPU T4 the error was not sticky; a valid index still worked next.

You named a GPU by number and that number is not in this process's device list.

Enum cudaErrorInvalidDevice
Code 101
Source https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29)

Strings a reader might paste

  • Runtime: invalid device ordinal
  • PyTorch adds a hint for this exact code and no other: after RuntimeError: CUDA error: invalid device ordinal it prints GPU device may be out of range, do you have enough GPUs?. Every other code gets the generic "search the docs" line. Source https://github.com/pytorch/pytorch/blob/main/c10/cuda/CUDAMiscFunctions.cpp , function get_cuda_error_help, checked 2026-08-29.

What the run shows, including the part that surprises people

From code/errors/invalid-device-ordinal/evidence/run-2026-09-01.txt, Tesla T4 (sm_75), driver 595.84, CUDA 12.6, on a node with exactly one GPU:

cudaGetDeviceCount                 -> 0 (cudaSuccess: no error)
device count: 1
cudaSetDevice(1)                   -> 101 (cudaErrorInvalidDevice: invalid device ordinal)
cudaSetDevice(0)                   -> 0 (cudaSuccess: no error)

The last line is the one worth keeping. Unlike the launch failures, 101 is not sticky: the process is entirely healthy afterwards and the very next call, with a legal index, succeeds. Nothing was corrupted, because nothing ran. That makes this error safe to handle in a loop, which is how a program should probe for the device it wants instead of assuming one.

The count of 1 on line two is the whole story of the bug. Ordinals are zero based, so the only valid index on this machine is 0, and 1 is one past the end.

Cause 1: an index past the end of the device list

cudaSetDevice(1) on a single-GPU machine, which is the measured case. It is also what happens when code written on a four-GPU workstation runs on a laptop.

int n = -1;
cudaGetDeviceCount(&n);      // check this, then bound your index by it
cudaSetDevice(1);

Fix: call cudaGetDeviceCount first and clamp or fail loudly against it. Never hardcode an ordinal in anything that will run on a second machine.

Cause 2: CUDA_VISIBLE_DEVICES renumbered the devices

This is the one that catches everybody once. The variable does not select which ordinal you use, it rebuilds the list. With CUDA_VISIBLE_DEVICES=2, the third physical GPU becomes device 0 inside your process, and asking for device 2 gives you 101 on a machine that visibly has four cards.

CUDA_VISIBLE_DEVICES=2 python train.py --gpu 2     # inside the process, only 0 exists

Fix: index from zero inside the process, always. The variable chooses the pool; your code chooses within it.

Cause 3: a stale index from a config or checkpoint

A launcher, a config file or a resumed checkpoint carrying cuda:3 from the machine it was written on.

Fix: derive the index at run time from the device count rather than storing it.

Confirm which one you have

Print the device count and the raw environment variable together, then compare them with the index you asked for:

echo "[$CUDA_VISIBLE_DEVICES]"

against the cudaGetDeviceCount line your program prints. If the count is smaller than your index, it is cause 1 or 2, and the bracketed echo tells you which: a non-empty value means the list was filtered and your index is being measured against the filtered list, not the machine.

Prevention

  • Treat the device count as the only source of truth about how many GPUs exist, and read it at start-up.
  • Log the ordinal and the device name together. "Using device 0 (Tesla T4)" in a log has settled more of these than any amount of reasoning about the environment.

Related errors

The lesson

Day 91 owns multi-GPU work and ordinal mapping in full. Day 9 covers what selecting a device costs, and the environment side lives at /setup/install-cuda. The full code table is on the error hub.


Written 2026-09-01. Transcript from code/errors/invalid-device-ordinal/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.