What is PTX in CUDA?
NVIDIA's virtual instruction set, which ships inside your binary and gets compiled to real machine code for the GPU that runs it.
PTX looks like assembly and is not the machine's. It is the stable intermediate language between nvcc's front end and ptxas, documented as its own ISA (https://docs.nvidia.com/cuda/parallel-thread-execution/index.html , checked 2026-08-30). The design reason it exists is time travel: embed PTX for a virtual architecture (compute_75) alongside or instead of SASS for real ones, and a GPU that did not exist when you shipped can still run your binary, because the driver JIT-compiles the PTX on first load. Which PTX and which SASS go into the fat binary is exactly what the -gencode flags decide.
The mistake to avoid is reading PTX as the compiler's final answer. It is an input to a second optimizing compiler, and register counts, instruction selection and scheduling are all decided after it. Day 46's fifth experiment nails this down mechanically: rebuild the same file with -Xptxas -O0 and the PTX changes by one header comment, no instructions, while the SASS diff is 428 KB. A flag that transformed the program completely was invisible at the PTX level, so conclusions read off PTX (register pressure especially) are conclusions about the wrong compiler's output.
What PTX is good for, day 46's side-by-side shows: it keeps your variable names and structure, so it is the readable bridge between your source and the SASS, and inline PTX (asm volatile) is how source code reaches instructions C++ has no spelling for.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75, captured 2026-09-01. Day 46 dumped PTX and SASS for the same three kernels (the .ptx, .sass and diff artifacts ship in the repo beside the code) and ran the flag experiment:
-Xptxas -O0 rebuild:
PTX diff: the ptxasOptions = -O0 header comment nvcc writes, nothing else
SASS diff: 428 KB
The run table for the three kernels (0.114, 0.113 and 0.157 ms, with the spill visible only in the SASS) is in the SASS entry; the division of labor between the two pages mirrors the division between the two formats.
Related terms
Where you meet this
- Day 46, PTX and SASS, the lesson that owns this term and ships the dumps.
- Day 3, installing CUDA, where the virtual-versus-real architecture split first appears.
- Day 17, registers and spills, whose ptxas reports describe the compilation PTX feeds.
Sources
- PTX ISA, chapter 1, for what PTX is and what it deliberately leaves to the next compiler: https://docs.nvidia.com/cuda/parallel-thread-execution/index.html (checked 2026-08-30)
- NVCC compiler driver documentation, for virtual and real architectures and the fat binary: https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.