← Glossary
CUDA glossaryTooling
CC 7.5

What is ptxas?

The assembler that turns PTX into SASS, decides how many registers each kernel gets, and reports when a kernel spills.

nvcc does not compile device code so much as hand it along. Your .cu file is split, the device half becomes PTX, and ptxas takes it from there: it assembles that virtual instruction set into real machine code for one named architecture and packs the result into a cubin. Along the way it makes the decision the rest of your kernel's performance hangs on, which is how many registers each thread gets. Nothing in your source states that number and no profiler is needed to learn it. It is a build-time fact.

The flag that prints it is -Xptxas -v, and -Xptxas is just nvcc's way of saying "pass this through". Ask for it and each kernel gets a short block naming the registers used per thread, the stack frame in bytes, the spill store and spill load counts, the shared memory it reserves and the constant bank usage per bank. Two of those fields are read wrong most often. The stack frame and the spill counts are different claims: a frame with no spills means the compiler routed something to local memory on purpose, while spill bytes mean it wanted registers and lost. And -O3 on your build line does not reach ptxas at all, because that flag sets the host optimisation level; device optimisation is -Xptxas -O3, which is a different compiler with its own switches.

Two more flags make the report do work rather than sit there. -Xptxas -warn-spills turns a spilling kernel into a build warning that names the function and the byte counts, and -Xptxas -Werror promotes that to a failure, so a regression stops the build instead of waiting to be noticed. None of this needs a card, which is why Compiler Explorer is enough: pick an NVCC entry rather than an NVRTC one, pin the version, and the ptxas info blocks land in the compiler output pane whether or not you ever press Execute.

Measured

Day 17 compiles four kernels out of one file, then asks the device what ptxas decided about each, through cudaFuncGetAttributes and cudaOccupancyMaxActiveBlocksPerMultiprocessor. Two independent views of one number are the only reason to trust either. Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), nvcc -O3 -arch=sm_75.

Kernel Registers Local memory Max threads per block Occupancy
decayNarrow 10 0 B 1024 100.0 %
decayWide 72 0 B 896 75.0 %
decayWideIndexed 69 256 B 896 75.0 %
decayWideBounded 64 48 B 256 100.0 %

Four kernels, one translation unit, four different allocations. Read the last column against the third: decayWideBounded is capped at 256 threads a block because its __launch_bounds__ says so, while decayWide is capped at 896, which is 65,536 registers divided by 72 rounded down to a whole number of warps. One honest gap: this run was captured without -Xptxas -v, so the report text itself is not on this page. That is the one artifact here you have to produce yourself, and it takes a rebuild and no hardware.

Related terms

Where you meet this

Sources

Byline

Author and reviewer are unassigned, and both have to be named before this entry leaves draft. The written and last-checked dates are set at that point.