How does nvcc gencode work?
The nvcc target flags that choose which real-GPU machine code and virtual-architecture PTX images a CUDA binary carries.
An sm_XX code target embeds SASS for that GPU family. A compute_XX code target embeds PTX that a compatible future driver may compile at run time. nvcc's compilation model distinguishes these real and virtual targets (https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html#gpu-compilation, checked 2026-09-01). A production fat binary usually carries several SASS images for known cards plus PTX as a forward-compatible fallback.
Do not infer the contents from the command line. Templates and separate translation units can create more than one image for a target. Inspect the linked executable with cuobjdump -lelf -lptx -all, then test the unsupported-card path because image loading can happen before the first launch.
Measured
On a Tesla T4 with driver 595.84 and CUDA 12.6, day 69 built one target with sm_75 SASS plus compute_90 PTX. cuobjdump listed 2 sm_75 cubins and 1 sm_90 PTX image. The normal T4 path passed at 4 resident blocks per SM, 160 blocks across 40 SMs. An sm_90-only binary failed at the occupancy query with no kernel image is available for execution on the device, exit 1.
Related terms
Where you meet this
- Day 69, CUDA portability, the measured fat-binary experiment.
- Day 67, CMake and CUDA, where target properties generate these flags.
- No kernel image is available, the runtime failure for a missing compatible image.
Sources
- NVIDIA nvcc manual, GPU compilation model: https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html#gpu-compilation (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came from the verification node on 2026-09-01.