CUDA ERROR
no kernel image is available for execution on the device
nvcc built for one architecture and the GPU is another. On a T4, an sm_80 binary returned CUDA error 209 at launch and success at the later sync.
Your binary contains no machine code and no forward-compatible PTX that can run on this GPU's compute capability, so the launch has nothing to load.
| Enum | cudaErrorNoKernelImageForDevice |
| Code | 209 |
| Source | https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29) |
Strings a reader might paste
- Runtime:
no kernel image is available for execution on the device - PyTorch warning, before any error, on 2.x:
NVIDIA GeForce RTX 5090 with CUDA capability sm_120 is not compatible with the current PyTorch installation.followed by the list of capabilities the wheel supports. Templateincompatible_device_warnin https://github.com/pytorch/pytorch/blob/main/torch/cuda/__init__.py (checked 2026-08-29). PyTorch main replaces it with a longer warning that startsFound GPU0 <name> which is of compute capability (CC) 12.0.and prints the exactpip install ... --index-urlline that would work (same file, checked 2026-08-29). - Sibling codes 200 and 300 both print
device kernel image is invalid: that is a corrupt or wrong-format image, where 209 is a missing one.
The measured run, and the part that fools your error check
We compiled a one-line kernel for sm_80 and ran it on this sm_75 T4. From code/errors/no-kernel-image/evidence/run-2026-09-01.txt, driver 595.84, CUDA 12.6:
== build ==
nvcc -std=c++17 -O3 -arch=sm_80 -o repro repro.cu # sm_80 binary, sm_75 card
== run ==
cudaMalloc -> 0 (cudaSuccess: no error)
peek after launch -> 209 (cudaErrorNoKernelImageForDevice: no kernel image is available for execution on the device)
cudaDeviceSynchronize -> 0 (cudaSuccess: no error)
The teaching point is the last line. 209 is reported at the launch, and the cudaDeviceSynchronize after it returns cudaSuccess, because no kernel ever ran and nothing is pending. A program that only checks after the sync sees a clean run that silently did nothing. cudaGetLastError() directly under the launch is the check that catches this one; day 6 is why every launch gets both lines.
Cause 1: the binary targets a newer architecture than your card
Exactly the measured case: -arch=sm_80 run on sm_75. Fix with a fat binary plus trailing PTX so newer cards can JIT:
nvcc -gencode arch=compute_75,code=sm_75 -gencode arch=compute_90,code=compute_90 ...
The trailing code=compute_90 embeds PTX. Note the direction: machine code does not run downward, and PTX only JITs upward. Also note the silent version of this bug: with no -arch at all, nvcc 13.x defaults to sm_75 (https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html#gpu-architecture-arch , "sm_75 is used as the default value", checked 2026-08-29), so an omitted flag builds a binary that lands on this page when run on an older card.
Cause 2: a prebuilt wheel with no kernels for a new card
The RTX 50 series (sm_120) against a PyTorch wheel built for sm_90 and below. Fix: install a wheel built against a CUDA release that includes your capability, using the index URL PyTorch's own warning prints. Evidence of the pattern: https://forums.developer.nvidia.com/t/rtx-5090-not-working-with-pytorch-and-stable-diffusion-sm-120-unsupported/338015 (checked 2026-08-29).
Cause 3: an architecture-specific target that does not carry forward
sm_90a code will not run on sm_100; the a suffix means no forward compatibility at all. wgmma requires sm_90a (https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#asynchronous-warpgroup-level-matrix-instructions-wgmma-mma , checked 2026-08-29). Fix: recompile per target; there is no other fix.
Confirm it
cuobjdump --list-elf ./a.out lists which architectures the binary embeds; compare against cudaDeviceGetAttribute with cudaDevAttrComputeCapabilityMajor and Minor. Honest note: on the node that produced the transcript above, cuobjdump was not installed (cuobjdump: command not found is in the evidence file), so that half of the check is documented here, not captured; the runtime half above is measured.
Prevention
- Always pass an explicit
-gencodelist; never rely on the default. - Check
cudaGetLastError()at the launch, since the sync will lie to you (measured above).
Related errors
invalid device function(98): a cousin where an image matched but the symbol is missing; on our 12.6 node the folklore repro for it did not fail at all, see that page.invalid configuration argument(9): also reported at launch, for illegal geometry rather than a missing image.nvcc fatal: Unsupported gpu architecture: the same mismatch caught at compile time instead.
The lesson
Day 46 covers what PTX and SASS are and why one JITs forward and the other does not; day 69 owns the full gencode portability story. The full code table is on the error hub; picking flags for a machine you have not met yet starts at /setup/install-cuda.
Written 2026-09-01. Transcript from code/errors/no-kernel-image/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.