What is CUDA JIT compilation?
The driver compiling embedded PTX into GPU machine code when a binary has no compatible prebuilt SASS image.
PTX is a virtual instruction set. When a CUDA binary has no suitable SASS image but does carry compatible PTX, the driver compiles that PTX for the installed GPU and normally caches the result. NVIDIA's fat-binary guide describes the JIT cache and CUDA_CACHE_DISABLE control (https://developer.nvidia.com/blog/cuda-pro-tip-understand-fat-binaries-jit-caching/, checked 2026-09-01).
JIT is a compatibility fallback, not a promise that every target moves forward forever. Architecture-specific targets such as Hopper's sm_90a include features that are not forward compatible; NVIDIA's Hopper compatibility guide distinguishes them from ordinary PTX fallback (https://docs.nvidia.com/cuda/hopper-compatibility-guide/, checked 2026-09-01). Ship matching SASS for predictable startup and PTX for compatible future cards.
Measured
On a Tesla T4 with driver 595.84 and CUDA 12.6, day 69 built a compute_75 PTX-only executable and ran it with CUDA_CACHE_DISABLE=1, forcing a real compile instead of a cache hit. It passed every gate and reported 4 active blocks per SM across 40 SMs, grid 160. The experiment did not time first-launch cost, so this page does not invent one.
Related terms
Where you meet this
- Day 69, CUDA portability, the cache-disabled PTX-only run.
- Day 9, timing, where first-use work is separated from steady state.
- No kernel image is available, the failure when no compatible image exists.
Sources
- NVIDIA, fat binaries and JIT caching: https://developer.nvidia.com/blog/cuda-pro-tip-understand-fat-binaries-jit-caching/ (checked 2026-09-01)
- NVIDIA Hopper compatibility guide, for forward-compatible PTX and architecture-specific exceptions: https://docs.nvidia.com/cuda/hopper-compatibility-guide/ (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came from the verification node on 2026-09-01.