CUDA errorsdraft
Reproduced 2026-09-01

CUDA ERROR

no kernel image is available for execution on the device

nvcc built for one architecture and the GPU is another. On a T4, an sm_80 binary returned CUDA error 209 at launch and success at the later sync.

Your binary contains no machine code and no forward-compatible PTX that can run on this GPU's compute capability, so the launch has nothing to load.

Enum cudaErrorNoKernelImageForDevice
Code 209
Source https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29)

Strings a reader might paste

  • Runtime: no kernel image is available for execution on the device
  • PyTorch warning, before any error, on 2.x: NVIDIA GeForce RTX 5090 with CUDA capability sm_120 is not compatible with the current PyTorch installation. followed by the list of capabilities the wheel supports. Template incompatible_device_warn in https://github.com/pytorch/pytorch/blob/main/torch/cuda/__init__.py (checked 2026-08-29). PyTorch main replaces it with a longer warning that starts Found GPU0 <name> which is of compute capability (CC) 12.0. and prints the exact pip install ... --index-url line that would work (same file, checked 2026-08-29).
  • Sibling codes 200 and 300 both print device kernel image is invalid: that is a corrupt or wrong-format image, where 209 is a missing one.

The measured run, and the part that fools your error check

We compiled a one-line kernel for sm_80 and ran it on this sm_75 T4. From code/errors/no-kernel-image/evidence/run-2026-09-01.txt, driver 595.84, CUDA 12.6:

== build ==
nvcc -std=c++17 -O3 -arch=sm_80 -o repro repro.cu   # sm_80 binary, sm_75 card

== run ==
cudaMalloc                         -> 0 (cudaSuccess: no error)
peek after launch                  -> 209 (cudaErrorNoKernelImageForDevice: no kernel image is available for execution on the device)
cudaDeviceSynchronize              -> 0 (cudaSuccess: no error)

The teaching point is the last line. 209 is reported at the launch, and the cudaDeviceSynchronize after it returns cudaSuccess, because no kernel ever ran and nothing is pending. A program that only checks after the sync sees a clean run that silently did nothing. cudaGetLastError() directly under the launch is the check that catches this one; day 6 is why every launch gets both lines.

Cause 1: the binary targets a newer architecture than your card

Exactly the measured case: -arch=sm_80 run on sm_75. Fix with a fat binary plus trailing PTX so newer cards can JIT:

nvcc -gencode arch=compute_75,code=sm_75 -gencode arch=compute_90,code=compute_90 ...

The trailing code=compute_90 embeds PTX. Note the direction: machine code does not run downward, and PTX only JITs upward. Also note the silent version of this bug: with no -arch at all, nvcc 13.x defaults to sm_75 (https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html#gpu-architecture-arch , "sm_75 is used as the default value", checked 2026-08-29), so an omitted flag builds a binary that lands on this page when run on an older card.

Cause 2: a prebuilt wheel with no kernels for a new card

The RTX 50 series (sm_120) against a PyTorch wheel built for sm_90 and below. Fix: install a wheel built against a CUDA release that includes your capability, using the index URL PyTorch's own warning prints. Evidence of the pattern: https://forums.developer.nvidia.com/t/rtx-5090-not-working-with-pytorch-and-stable-diffusion-sm-120-unsupported/338015 (checked 2026-08-29).

Cause 3: an architecture-specific target that does not carry forward

sm_90a code will not run on sm_100; the a suffix means no forward compatibility at all. wgmma requires sm_90a (https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#asynchronous-warpgroup-level-matrix-instructions-wgmma-mma , checked 2026-08-29). Fix: recompile per target; there is no other fix.

Confirm it

cuobjdump --list-elf ./a.out lists which architectures the binary embeds; compare against cudaDeviceGetAttribute with cudaDevAttrComputeCapabilityMajor and Minor. Honest note: on the node that produced the transcript above, cuobjdump was not installed (cuobjdump: command not found is in the evidence file), so that half of the check is documented here, not captured; the runtime half above is measured.

Prevention

  • Always pass an explicit -gencode list; never rely on the default.
  • Check cudaGetLastError() at the launch, since the sync will lie to you (measured above).

Related errors

The lesson

Day 46 covers what PTX and SASS are and why one JITs forward and the other does not; day 69 owns the full gencode portability story. The full code table is on the error hub; picking flags for a machine you have not met yet starts at /setup/install-cuda.


Written 2026-09-01. Transcript from code/errors/no-kernel-image/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.