CUDA / GLOSSARY
GPU terms, explained
120terms with examples and sources
01
22 terms
Execution model
22
- 01__syncthreads()A barrier that every thread in a block must reach before any thread passes it, and which also makes shared-memory writes visible to the rest of the block.CC 7.5
- 02built-in index variables`threadIdx`, `blockIdx`, `blockDim` and `gridDim` are the four values every thread reads to work out which piece of the data it owns.CC 7.5
- 03cooperative groupsA C++ API for naming the group you want to synchronize, from a 32-lane tile up to the whole grid.CC 7.5
- 04CUDA eventA marker you record into a stream, used either to time GPU work or to make one stream wait on another.CC 7.5
- 05CUDA graphA captured set of kernels and copies with their dependencies, launched as one object so you pay the launch cost once.CC 7.5
- 06CUDA streamAn ordered queue of GPU work, where two streams may overlap and one stream may not.CC 7.5
- 07cudaDeviceSynchronize()Blocks the calling host thread until every previously launched piece of GPU work has finished.CC 7.5
- 08dynamic parallelismA kernel launching another kernel from the device, which CUDA 12 rebuilt as CDP2.CC 7.5
- 09execution configurationThe `<<<blocks, threads, sharedBytes, stream>>>` syntax between a kernel's name and its arguments.CC 7.5
- 10execution space specifiers`__global__` marks a kernel the host launches, `__device__` marks a function only device code can call, and `__host__` marks ordinary CPU code.CC 7.5
- 11global thread indexThe single number that identifies a thread across the whole grid, usually `blockIdx.x * blockDim.x + threadIdx.x`.CC 7.5
- 12gridAll the blocks one kernel launch creates, in one, two or three dimensions.CC 7.5
- 13grid-stride loopA loop that lets a fixed-size grid cover any N by having each thread step forward by the total thread count.CC 7.5
- 14independent thread schedulingSince Volta, each thread has its own program counter, so lanes of a warp can be at different instructions and the old lockstep assumptions break.CC 7.0
- 15kernelA function you mark `__global__` and launch across a grid of threads, which runs on the GPU while the CPU keeps going.CC 7.5
- 16laneA thread's position inside its warp, 0 to 31, which is what shuffle and vote instructions address.CC 7.5
- 17SIMTSingle instruction, multiple threads: the hardware issues one instruction to 32 threads that each keep their own registers and their own program counter.CC any
- 18threadOne instance of your kernel body, with its own registers and its own index, scheduled as part of a group of 32.CC 7.5
- 19thread blockA group of threads that lands on one SM, shares that SM's shared memory, and can synchronize with `__syncthreads()`.CC 7.5
- 20warpThe 32 threads an SM issues together, sharing one instruction stream and one set of memory requests.CC 7.5
- 21warp divergenceWhat happens when lanes of one warp take different branches, so the hardware runs both paths with some lanes masked off.CC 7.5
- 22warp shuffleInstructions that let lanes of one warp read each other's registers directly, with no shared memory and no barrier.CC 7.5
02
22 terms
Memory
22
- 01bank conflictTwo lanes of a warp hitting different addresses in the same bank, which serializes their accesses.CC 7.5
- 02constant memoryA 64 KB read-only region of device memory with its own cache, fastest exactly when every lane of a warp reads the same address.CC 7.5
- 03cudaMemcpyThe blocking copy between host and device, whose direction argument is the first thing beginners get wrong.CC 7.5
- 04global memoryThe GPU's main DRAM, visible to every thread, and the slowest thing a kernel touches.CC 7.5
- 05GPU RAMThe DRAM attached to a GPU, usually GDDR on consumer cards or stacked HBM on datacenter accelerators.CC 7.5
- 06L1 and L2 cacheL1 is per SM and shares its silicon with shared memory; L2 is one cache for the whole device, and every access to DRAM goes through it.CC 7.5
- 07local memoryPer-thread storage that lives off chip in device DRAM despite its name, holding whatever the compiler cannot keep in registers.CC 7.5
- 08memory bankOne of 32 independent slices of shared memory, each serving 4 bytes per cycle, addressed by `(byteAddress / 4) % 32`.CC 7.5
- 09memory coalescingThe hardware merging a warp's 32 addresses into the fewest memory transactions that cover them: cheap when the addresses are consecutive, expensive when they are not.CC 7.5
- 10memory hierarchyThe stack of storage a kernel can reach, from registers through shared memory and caches out to global memory, each one bigger and slower than the last.CC 7.5
- 11memory transactionOne request the memory system serves for a warp, covering a fixed span of bytes whether or not the warp wanted all of them.CC 7.5
- 12padding and swizzlingTwo fixes for bank conflicts: add a column so rows land on different banks, or XOR the index so the mapping rotates.CC 7.5
- 13page migrationThe unified-memory driver moving a virtual-memory page to the processor that touches it, either after a fault or ahead of use through prefetch.CC depends on managed-memory attributes; lesson requires 7.5
- 14pinned memoryHost memory the OS cannot page out, which lets the GPU copy from it by DMA and lets copies overlap kernels.CC 7.5
- 15register spillingThe compiler running out of registers and pushing values into local memory, so every later use of them costs a trip to device DRAM.CC 7.5
- 16registersThe fastest storage on the chip, private to one thread, handed out by the compiler and capped by a fixed pool each SM divides among its resident threads.CC 7.5
- 17sector and cache lineA sector is the 32-byte unit the memory system actually fetches, and a 128-byte cache line is four of them.CC 7.5
- 18shared memoryFast per-block scratch memory on the SM, which you fill yourself and which vanishes when the block ends.CC 7.5
- 19stream-ordered allocationAllocating and freeing device memory inside a stream, from a pool, so the call does not synchronize the whole device.CC 7.5
- 20tilingLoading a block-sized piece of the input into shared memory once, then reusing it from there instead of going back to global memory.CC 7.5
- 21unified memoryOne pointer valid on both the host and the device, with the driver moving pages between them when the other side touches them.CC 7.5
- 22vectorized loadReading 8 or 16 bytes per lane with float2 or float4, so one warp covers more bytes per instruction.CC 7.5
03
17 terms
Performance
17
- 01arithmetic intensityFlops per byte moved, the one number that says whether a kernel is limited by math or by memory.CC any
- 02atomic contentionMany threads hitting the same address with atomics, which serializes them no matter how many threads you launched.CC 7.5
- 03compute-boundA kernel whose math pipelines are the limit, so better memory access buys nothing.CC any
- 04kernel fusionMerging several kernels into one so intermediate results stay in registers instead of round-tripping through global memory.CC 7.5
- 05latency hidingKeeping enough warps in flight that the SM always has one ready to issue while the others wait on memory.CC 7.5
- 06launch overheadThe fixed cost of getting a kernel started, which dominates when the kernel itself is short.CC 7.5
- 07Little's lawConcurrency equals latency times throughput, which is why you need many memory requests in flight to saturate a GPU's bandwidth.CC 7.5
- 08loop unrollingThe compiler replacing a loop with repeated copies of its body, which removes branch work and exposes independent instructions.CC 7.5
- 09memory bandwidthBytes per second between the SMs and DRAM, quoted as a theoretical peak on the box and measured as an effective figure by a real kernel.CC 7.5
- 10memory-boundA kernel that spends its time waiting for data, so making the math faster changes nothing.CC any
- 11occupancyThe fraction of an SM's warp slots your kernel actually fills, which is a means of hiding latency and not a goal in itself.CC 7.5
- 12online softmaxA one-pass softmax recurrence that rescales its running sum whenever the maximum increases, allowing normalization without storing all logits first.CC 7.5
- 13privatizationGiving each block its own copy of a contended structure in shared memory, then merging the copies once at the end.CC 7.5
- 14register pressureThe tension between giving each thread more registers and keeping enough threads resident to hide latency.CC 7.5
- 15roofline modelA plot of achievable throughput against arithmetic intensity, with a sloped memory limit and a flat compute limit, that tells you which one you are hitting.CC 7.5
- 16speed of lightNsight Compute's headline percentage: how close a kernel gets to the hardware's maximum for compute and for memory.CC 7.5
- 17warp stall reasonsThe named categories Nsight Compute reports for why a warp could not issue, such as long scoreboard, barrier or MIO throttle.CC 7.5
04
15 terms
Tooling
15
- 01compute-sanitizerA suite of runtime correctness tools for CUDA memory accesses, shared-memory races, synchronization and uninitialised data.CC 7.5
- 02CUDA driver APIThe lower-level cu* interface in libcuda, shipped with the driver rather than the toolkit, which is why its version differs from nvcc's.CC 7.5
- 03CUDA error checkingEvery runtime call returns a status and every kernel launch fails silently unless you ask, which is what `cudaGetLastError` and a check macro are for.CC 7.5
- 04CUDA runtime APIThe cuda* calls in cudart that almost all CUDA code uses, layered on top of the driver API.CC 7.5
- 05cuda-gdbThe CUDA-aware debugger that can stop inside a kernel, select a GPU thread and inspect its device-side state.CC 7.5
- 06CUPTIThe profiling interface underneath Nsight and the PyTorch profiler, which is where the counter permission requirement comes from.CC 7.5
- 07lazy module loadingThe driver deferring each kernel's load until its first launch, which moves a cost most benchmarks then measure by accident.CC 7.5
- 08Nsight ComputeThe kernel profiler, which replays one kernel and reports its memory chart, its stall reasons and how close it got to the hardware limits.CC 7.5
- 09Nsight SystemsThe timeline profiler, which shows what the CPU, the copies and the kernels were each doing and when.CC 7.5
- 10nvccThe CUDA compiler driver, which splits your file into host code for the system compiler and device code for ptxas.CC 7.5
- 11NVTXAnnotations you add to your own code so a profiler timeline shows your phase names instead of anonymous bars.CC 7.5
- 12PTXNVIDIA's virtual instruction set, which ships inside your binary and gets compiled to real machine code for the GPU that runs it.CC 7.5
- 13ptxasThe assembler that turns PTX into SASS, decides how many registers each kernel gets, and reports when a kernel spills.CC 7.5
- 14PyTorch profilerPyTorch's operator-level profiler, which records CPU and CUDA activity, input shapes and traces that connect framework operations to device kernels.CC any CUDA GPU supported by the installed PyTorch build
- 15SASSThe actual machine code a GPU executes, which you read with cuobjdump or nvdisasm and which is where the truth lives.CC 7.5
05
12 terms
Hardware
12
- 01compute capabilityThe version number that says which features a GPU has, written 7.5 in docs and sm_75 on the nvcc command line.CC 7.5
- 02CUDA coreA lane of the SM's arithmetic pipeline, not a core in the CPU sense, and the number on the box is a count of these.CC any
- 03GPU architecture generationsTuring, Ampere, Ada, Hopper and Blackwell, each mapping to a range of compute capabilities and a set of features.CC 7.5
- 04green contextA CUDA execution context provisioned with a selected group of streaming multiprocessors, restricting its work to that SM partition.CC the measured binary targets 7.5; API availability depends on toolkit and driver
- 05host and deviceHost is the CPU and its memory, device is the GPU and its memory, and nothing crosses between them without a copy or a managed pointer.CC 7.5
- 06MPS and MIGTwo GPU-sharing mechanisms: MPS coordinates CUDA processes through a server, while MIG partitions supported datacenter GPUs into isolated device instances.CC MIG begins with supported datacenter Ampere GPUs; MPS feature floors vary
- 07PCIeThe bus between host memory and the GPU, roughly an order of magnitude slower than the GPU's own memory.CC 7.5
- 08register fileThe fixed pool of registers on each SM, partitioned among every thread resident there, which is what caps occupancy in most real kernels.CC 7.5
- 09resident warps and blocks per SMThe hard caps on how many warps and blocks one SM can hold at once, which set the ceiling occupancy can reach.CC 7.5
- 10streaming multiprocessorThe unit a GPU is actually made of: it owns a slice of registers and shared memory, holds many blocks at once, and issues warps from whichever one is ready.CC any
- 11tensor coreA specialized arithmetic unit that performs a small matrix multiply-accumulate cooperatively for a warp, at throughput ordinary scalar CUDA-core instructions cannot match.CC 7.0; lesson path requires 7.5
- 12warp schedulerThe part of the SM that picks, every cycle, which resident warp gets to issue an instruction.CC 7.5
06
15 terms
Precision
15
- 012:4 structured sparsityKeeping exactly two values in every consecutive group of four so supported sparse tensor cores can skip a predictable half of a matrix.CC 8.0 for 2:4 sparse tensor-core acceleration
- 02accumulator typeThe data type used for a running sum or matrix output fragment, chosen independently from input storage to control rounding, range and hardware behavior.CC 7.5
- 03BF16A 16-bit float with FP32's eight exponent bits and seven stored mantissa bits, preserving range while giving up precision near ordinary values.CC 8.0 for tensor-core mma
- 04fast mathThe -use_fast_math flag, which swaps precise math functions for faster approximations and turns denormal support off.CC 7.5
- 05floating point determinismRepeatable floating-point results require a fixed input, reduction shape, addition order and compiled instruction sequence.CC 7.5
- 06FMAFused multiply-add: one instruction that computes a*b+c with a single rounding, which is both faster and more accurate than doing it in two steps.CC 7.5
- 07FP16A 16-bit floating-point format with five exponent bits and ten stored mantissa bits, giving compact storage but a maximum finite value of 65,504.CC 7.5
- 08FP32 and FP64The 32-bit single-precision and 64-bit double-precision IEEE formats, used respectively for ordinary GPU arithmetic and higher-accuracy references or scientific work.CC any
- 09FP8A family of eight-bit floating-point formats that trade range against precision, commonly using E4M3 for activations and E5M2 where wider range matters.CC 8.9 for tensor-core mma
- 10INT8 quantizationRepresenting real values as signed eight-bit integers plus scale and optional zero point, then accumulating integer products and rescaling the output.CC 7.5 for INT8 tensor cores; measured kernel uses CUDA cores
- 11mixed precisionComputing with narrow input or storage formats while retaining a wider accumulator, so bandwidth and arithmetic throughput improve without paying narrow-format error at every sum.CC 7.5
- 12mma.syncA PTX warp-level matrix instruction whose shape, layouts and data types determine exactly which registers each lane supplies and receives.CC 7.5 for m16n8k8 FP16
- 13MX block scalingAssigning one scale to a small block of low-precision values so each block can use the element format's limited range more closely.CC format-dependent; the day 98 T4 measurement is a host storage census
- 14TF32A tensor-core input format with FP32's range and at least ten bits of precision, used by Ampere-or-newer matrix instructions rather than ordinary CUDA-core arithmetic.CC 8.0
- 15WMMACUDA's warp-level matrix API, where 32 threads cooperatively load opaque fragments, execute a matrix multiply-accumulate, and store the resulting tile.CC 7.0; lesson path uses sm_75
07
15 terms
Libraries
15
- 01CCCLThe CUDA C++ Core Libraries, one repo holding Thrust, CUB and libcu++, versioned and shipped together.CC 7.5
- 02CUBThrust's engine, exposed at block, warp and device scope so you can drop a tuned primitive inside your own kernel.CC 7.5
- 03cuBLASNVIDIA's dense linear algebra library, providing architecture-tuned matrix operations whose data type, compute type and math mode define the numerical contract.CC any supported CUDA GPU
- 04cuBLASLtA descriptor-based matrix multiplication interface that exposes algorithm heuristics, workspace choices, layouts and fused epilogues beyond traditional cuBLAS GEMM calls.CC any supported CUDA GPU
- 05CUDA math librariesToolkit libraries that provide tuned random generation, transforms, sparse algebra and dense solvers, replacing custom kernels while exposing plans, state and scratch-buffer costs.CC any supported CUDA GPU
- 06cuDNNNVIDIA's deep-learning primitive library, whose graph API turns described operations into a selected engine, execution plan, workspace and bound tensor addresses.CC depends on operation and engine; measured FP32 graph ran at 7.5
- 07cuSPARSENVIDIA's sparse linear algebra library, providing descriptor-based operations such as CSR sparse matrix-vector multiply with algorithm and temporary-storage choices.CC any supported CUDA GPU
- 08CuTeCUTLASS's layout algebra and C++ type system, where a shape and stride define a coordinate-to-index function that can be composed and divided into tiled mappings.CC layout algebra is host-capable; measured device gather ran at 7.5
- 09CUTLASSNVIDIA's header-only CUDA C++ template library for composing architecture-specific matrix multiplication and related kernels from explicit data, tile, pipeline and epilogue choices.CC depends on the selected kernel; measured sm_75 path runs at 7.5
- 10gencodeThe nvcc target flags that choose which real-GPU machine code and virtual-architecture PTX images a CUDA binary carries.CC 7.5
- 11JIT compilationThe driver compiling embedded PTX into GPU machine code when a binary has no compatible prebuilt SASS image.CC 7.5
- 12libcu++The CUDA standard library: cuda::std types that work on both sides, plus device primitives like cuda::atomic_ref, cuda::barrier and cuda::pipeline.CC 7.5
- 13PyTorch custom operatorA user-defined operation registered with PyTorch's dispatcher, with explicit device implementations and optional fake and autograd registrations.CC set by the registered kernel; day 87 requires 7.5
- 14separate compilationCompiling device code in multiple translation units and joining it with a device-link step before the final host link.CC 7.5
- 15ThrustAn STL-shaped library of parallel algorithms over device data, where thrust::reduce replaces a kernel you would otherwise write.CC 7.5
08
2 terms
Ecosystem
2