← Glossary
CUDA glossaryLibraries
CC 7.5

What is CUDA JIT compilation?

The driver compiling embedded PTX into GPU machine code when a binary has no compatible prebuilt SASS image.

PTX is a virtual instruction set. When a CUDA binary has no suitable SASS image but does carry compatible PTX, the driver compiles that PTX for the installed GPU and normally caches the result. NVIDIA's fat-binary guide describes the JIT cache and CUDA_CACHE_DISABLE control (https://developer.nvidia.com/blog/cuda-pro-tip-understand-fat-binaries-jit-caching/, checked 2026-09-01).

JIT is a compatibility fallback, not a promise that every target moves forward forever. Architecture-specific targets such as Hopper's sm_90a include features that are not forward compatible; NVIDIA's Hopper compatibility guide distinguishes them from ordinary PTX fallback (https://docs.nvidia.com/cuda/hopper-compatibility-guide/, checked 2026-09-01). Ship matching SASS for predictable startup and PTX for compatible future cards.

Measured

On a Tesla T4 with driver 595.84 and CUDA 12.6, day 69 built a compute_75 PTX-only executable and ran it with CUDA_CACHE_DISABLE=1, forcing a real compile instead of a cache hit. It passed every gate and reported 4 active blocks per SM across 40 SMs, grid 160. The experiment did not time first-launch cost, so this page does not invent one.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came from the verification node on 2026-09-01.