What is ptxas?
The assembler that turns PTX into SASS, decides how many registers each kernel gets, and reports when a kernel spills.
nvcc does not compile device code so much as hand it along. Your .cu file is split, the device half becomes PTX, and ptxas takes it from there: it assembles that virtual instruction set into real machine code for one named architecture and packs the result into a cubin. Along the way it makes the decision the rest of your kernel's performance hangs on, which is how many registers each thread gets. Nothing in your source states that number and no profiler is needed to learn it. It is a build-time fact.
The flag that prints it is -Xptxas -v, and -Xptxas is just nvcc's way of saying "pass this through". Ask for it and each kernel gets a short block naming the registers used per thread, the stack frame in bytes, the spill store and spill load counts, the shared memory it reserves and the constant bank usage per bank. Two of those fields are read wrong most often. The stack frame and the spill counts are different claims: a frame with no spills means the compiler routed something to local memory on purpose, while spill bytes mean it wanted registers and lost. And -O3 on your build line does not reach ptxas at all, because that flag sets the host optimisation level; device optimisation is -Xptxas -O3, which is a different compiler with its own switches.
Two more flags make the report do work rather than sit there. -Xptxas -warn-spills turns a spilling kernel into a build warning that names the function and the byte counts, and -Xptxas -Werror promotes that to a failure, so a regression stops the build instead of waiting to be noticed. None of this needs a card, which is why Compiler Explorer is enough: pick an NVCC entry rather than an NVRTC one, pin the version, and the ptxas info blocks land in the compiler output pane whether or not you ever press Execute.
Measured
Day 17 compiles four kernels out of one file, then asks the device what ptxas decided about each, through cudaFuncGetAttributes and cudaOccupancyMaxActiveBlocksPerMultiprocessor. Two independent views of one number are the only reason to trust either. Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), nvcc -O3 -arch=sm_75.
| Kernel | Registers | Local memory | Max threads per block | Occupancy |
|---|---|---|---|---|
decayNarrow |
10 | 0 B | 1024 | 100.0 % |
decayWide |
72 | 0 B | 896 | 75.0 % |
decayWideIndexed |
69 | 256 B | 896 | 75.0 % |
decayWideBounded |
64 | 48 B | 256 | 100.0 % |
Four kernels, one translation unit, four different allocations. Read the last column against the third: decayWideBounded is capped at 256 threads a block because its __launch_bounds__ says so, while decayWide is capped at 896, which is 65,536 registers divided by 72 rounded down to a whole number of warps. One honest gap: this run was captured without -Xptxas -v, so the report text itself is not on this page. That is the one artifact here you have to produce yourself, and it takes a rebuild and no hardware.
Related terms
Where you meet this
- Day 17, registers, local memory and spills, the lesson that owns this term and turns the report into occupancy arithmetic.
- How to set up CUDA, for the build line the flag hangs off.
- Learn CUDA without a GPU, since the report is a build artifact and a hosted compiler will print it.
too many resources requested for launch, the run-time end of a decision ptxas made at build time.
Sources
- NVIDIA CUDA Compiler Driver NVCC, 8.4 "Printing Code Generation Statistics", plus the
-Xptxasand-maxrregcountoptions: https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html (checked 2026-08-30) - CUDA Programming Guide, 5.4.3.2 "Launch Bounds", for the register ceiling ptxas derives from a bound: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/cpp-language-extensions.html (checked 2026-08-30)
- Compiler Explorer's CUDA compiler list, showing which entries are real NVCC and therefore print a ptxas report: https://godbolt.org/api/compilers/cuda?fields=id,name,supportsExecute (checked 2026-08-29)
Byline
Author and reviewer are unassigned, and both have to be named before this entry leaves draft. The written and last-checked dates are set at that point.