← Glossary
CUDA glossaryTooling
CC 7.5

What is SASS in CUDA?

The actual machine code a GPU executes, which you read with cuobjdump or nvdisasm and which is where the truth lives.

Your C++ becomes PTX, and PTX becomes SASS, the per-architecture machine code, when ptxas compiles it. Every question about what the compiler really did, whether the loop unrolled, whether the multiply-add fused, whether a value spilled, has its answer here and only here, because both earlier forms are inputs to further optimization. cuobjdump -sass binary dumps it, nvdisasm adds control-flow detail, and both ship with the toolkit (https://docs.nvidia.com/cuda/cuda-binary-utilities/index.html , checked 2026-08-30).

You do not need to write SASS or even read it fluently; you need to find things in it, and two hunts cover most real use. Hunt one, control flow: a surviving loop is a backward BRA, so counting back edges answers "did my loop unroll". Hunt two, spills: STL and LDL are local-memory stores and loads, so their presence answers "did my array stay in registers". Both hunts are greps, not analysis, and day 46 does each against a kernel built to contain the answer.

The same day's fifth experiment is the sharpest argument for reading SASS rather than PTX: rebuilding with -Xptxas -O0 changed the PTX by exactly one header comment and no instruction, while the SASS diff was 428 KB. The flag reaches only the second compiler, so the intermediate form cannot show you what it did.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75, disassembled with the same toolkit's cuobjdump and nvdisasm, captured 2026-09-01. Day 46 compiled three variants of one 64-tap decaying fold and counted what the SASS holds:

kernel back edges FFMA STL, LDL regs local B ms
decayRolled 2 29 0, 0 42 0 0.114
decayUnrolled 0 63 0, 0 62 0 0.113
decayStaged 0 126 11, 11 64 48 0.157

Each row is a found thing, not a timing: the rolled loop keeps two backward branches (a loop and its remainder), the unrolled one keeps none, and only the staged variant carries the spill instructions, matching its 48 bytes of local memory from ptxas and its measured slowdown. The 29 FFMAs in the rolled kernel are themselves a finding: the compiler partially unrolled the "rolled" loop on its own, which only the disassembly could reveal.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.