CUDA / GLOSSARY

GPU terms, explained

120terms with examples and sources
01

22 terms

Execution model

22
  1. 01
    __syncthreads()A barrier that every thread in a block must reach before any thread passes it, and which also makes shared-memory writes visible to the rest of the block.
    CC 7.5
  2. 02
    built-in index variables`threadIdx`, `blockIdx`, `blockDim` and `gridDim` are the four values every thread reads to work out which piece of the data it owns.
    CC 7.5
  3. 03
    cooperative groupsA C++ API for naming the group you want to synchronize, from a 32-lane tile up to the whole grid.
    CC 7.5
  4. 04
    CUDA eventA marker you record into a stream, used either to time GPU work or to make one stream wait on another.
    CC 7.5
  5. 05
    CUDA graphA captured set of kernels and copies with their dependencies, launched as one object so you pay the launch cost once.
    CC 7.5
  6. 06
    CUDA streamAn ordered queue of GPU work, where two streams may overlap and one stream may not.
    CC 7.5
  7. 07
    cudaDeviceSynchronize()Blocks the calling host thread until every previously launched piece of GPU work has finished.
    CC 7.5
  8. 08
    dynamic parallelismA kernel launching another kernel from the device, which CUDA 12 rebuilt as CDP2.
    CC 7.5
  9. 09
    execution configurationThe `<<<blocks, threads, sharedBytes, stream>>>` syntax between a kernel's name and its arguments.
    CC 7.5
  10. 10
    execution space specifiers`__global__` marks a kernel the host launches, `__device__` marks a function only device code can call, and `__host__` marks ordinary CPU code.
    CC 7.5
  11. 11
    global thread indexThe single number that identifies a thread across the whole grid, usually `blockIdx.x * blockDim.x + threadIdx.x`.
    CC 7.5
  12. 12
    gridAll the blocks one kernel launch creates, in one, two or three dimensions.
    CC 7.5
  13. 13
    grid-stride loopA loop that lets a fixed-size grid cover any N by having each thread step forward by the total thread count.
    CC 7.5
  14. 14
    independent thread schedulingSince Volta, each thread has its own program counter, so lanes of a warp can be at different instructions and the old lockstep assumptions break.
    CC 7.0
  15. 15
    kernelA function you mark `__global__` and launch across a grid of threads, which runs on the GPU while the CPU keeps going.
    CC 7.5
  16. 16
    laneA thread's position inside its warp, 0 to 31, which is what shuffle and vote instructions address.
    CC 7.5
  17. 17
    SIMTSingle instruction, multiple threads: the hardware issues one instruction to 32 threads that each keep their own registers and their own program counter.
    CC any
  18. 18
    threadOne instance of your kernel body, with its own registers and its own index, scheduled as part of a group of 32.
    CC 7.5
  19. 19
    thread blockA group of threads that lands on one SM, shares that SM's shared memory, and can synchronize with `__syncthreads()`.
    CC 7.5
  20. 20
    warpThe 32 threads an SM issues together, sharing one instruction stream and one set of memory requests.
    CC 7.5
  21. 21
    warp divergenceWhat happens when lanes of one warp take different branches, so the hardware runs both paths with some lanes masked off.
    CC 7.5
  22. 22
    warp shuffleInstructions that let lanes of one warp read each other's registers directly, with no shared memory and no barrier.
    CC 7.5
02

22 terms

Memory

22
  1. 01
    bank conflictTwo lanes of a warp hitting different addresses in the same bank, which serializes their accesses.
    CC 7.5
  2. 02
    constant memoryA 64 KB read-only region of device memory with its own cache, fastest exactly when every lane of a warp reads the same address.
    CC 7.5
  3. 03
    cudaMemcpyThe blocking copy between host and device, whose direction argument is the first thing beginners get wrong.
    CC 7.5
  4. 04
    global memoryThe GPU's main DRAM, visible to every thread, and the slowest thing a kernel touches.
    CC 7.5
  5. 05
    GPU RAMThe DRAM attached to a GPU, usually GDDR on consumer cards or stacked HBM on datacenter accelerators.
    CC 7.5
  6. 06
    L1 and L2 cacheL1 is per SM and shares its silicon with shared memory; L2 is one cache for the whole device, and every access to DRAM goes through it.
    CC 7.5
  7. 07
    local memoryPer-thread storage that lives off chip in device DRAM despite its name, holding whatever the compiler cannot keep in registers.
    CC 7.5
  8. 08
    memory bankOne of 32 independent slices of shared memory, each serving 4 bytes per cycle, addressed by `(byteAddress / 4) % 32`.
    CC 7.5
  9. 09
    memory coalescingThe hardware merging a warp's 32 addresses into the fewest memory transactions that cover them: cheap when the addresses are consecutive, expensive when they are not.
    CC 7.5
  10. 10
    memory hierarchyThe stack of storage a kernel can reach, from registers through shared memory and caches out to global memory, each one bigger and slower than the last.
    CC 7.5
  11. 11
    memory transactionOne request the memory system serves for a warp, covering a fixed span of bytes whether or not the warp wanted all of them.
    CC 7.5
  12. 12
    padding and swizzlingTwo fixes for bank conflicts: add a column so rows land on different banks, or XOR the index so the mapping rotates.
    CC 7.5
  13. 13
    page migrationThe unified-memory driver moving a virtual-memory page to the processor that touches it, either after a fault or ahead of use through prefetch.
    CC depends on managed-memory attributes; lesson requires 7.5
  14. 14
    pinned memoryHost memory the OS cannot page out, which lets the GPU copy from it by DMA and lets copies overlap kernels.
    CC 7.5
  15. 15
    register spillingThe compiler running out of registers and pushing values into local memory, so every later use of them costs a trip to device DRAM.
    CC 7.5
  16. 16
    registersThe fastest storage on the chip, private to one thread, handed out by the compiler and capped by a fixed pool each SM divides among its resident threads.
    CC 7.5
  17. 17
    sector and cache lineA sector is the 32-byte unit the memory system actually fetches, and a 128-byte cache line is four of them.
    CC 7.5
  18. 18
    shared memoryFast per-block scratch memory on the SM, which you fill yourself and which vanishes when the block ends.
    CC 7.5
  19. 19
    stream-ordered allocationAllocating and freeing device memory inside a stream, from a pool, so the call does not synchronize the whole device.
    CC 7.5
  20. 20
    tilingLoading a block-sized piece of the input into shared memory once, then reusing it from there instead of going back to global memory.
    CC 7.5
  21. 21
    unified memoryOne pointer valid on both the host and the device, with the driver moving pages between them when the other side touches them.
    CC 7.5
  22. 22
    vectorized loadReading 8 or 16 bytes per lane with float2 or float4, so one warp covers more bytes per instruction.
    CC 7.5
03

17 terms

Performance

17
  1. 01
    arithmetic intensityFlops per byte moved, the one number that says whether a kernel is limited by math or by memory.
    CC any
  2. 02
    atomic contentionMany threads hitting the same address with atomics, which serializes them no matter how many threads you launched.
    CC 7.5
  3. 03
    compute-boundA kernel whose math pipelines are the limit, so better memory access buys nothing.
    CC any
  4. 04
    kernel fusionMerging several kernels into one so intermediate results stay in registers instead of round-tripping through global memory.
    CC 7.5
  5. 05
    latency hidingKeeping enough warps in flight that the SM always has one ready to issue while the others wait on memory.
    CC 7.5
  6. 06
    launch overheadThe fixed cost of getting a kernel started, which dominates when the kernel itself is short.
    CC 7.5
  7. 07
    Little's lawConcurrency equals latency times throughput, which is why you need many memory requests in flight to saturate a GPU's bandwidth.
    CC 7.5
  8. 08
    loop unrollingThe compiler replacing a loop with repeated copies of its body, which removes branch work and exposes independent instructions.
    CC 7.5
  9. 09
    memory bandwidthBytes per second between the SMs and DRAM, quoted as a theoretical peak on the box and measured as an effective figure by a real kernel.
    CC 7.5
  10. 10
    memory-boundA kernel that spends its time waiting for data, so making the math faster changes nothing.
    CC any
  11. 11
    occupancyThe fraction of an SM's warp slots your kernel actually fills, which is a means of hiding latency and not a goal in itself.
    CC 7.5
  12. 12
    online softmaxA one-pass softmax recurrence that rescales its running sum whenever the maximum increases, allowing normalization without storing all logits first.
    CC 7.5
  13. 13
    privatizationGiving each block its own copy of a contended structure in shared memory, then merging the copies once at the end.
    CC 7.5
  14. 14
    register pressureThe tension between giving each thread more registers and keeping enough threads resident to hide latency.
    CC 7.5
  15. 15
    roofline modelA plot of achievable throughput against arithmetic intensity, with a sloped memory limit and a flat compute limit, that tells you which one you are hitting.
    CC 7.5
  16. 16
    speed of lightNsight Compute's headline percentage: how close a kernel gets to the hardware's maximum for compute and for memory.
    CC 7.5
  17. 17
    warp stall reasonsThe named categories Nsight Compute reports for why a warp could not issue, such as long scoreboard, barrier or MIO throttle.
    CC 7.5
04

15 terms

Tooling

15
  1. 01
    compute-sanitizerA suite of runtime correctness tools for CUDA memory accesses, shared-memory races, synchronization and uninitialised data.
    CC 7.5
  2. 02
    CUDA driver APIThe lower-level cu* interface in libcuda, shipped with the driver rather than the toolkit, which is why its version differs from nvcc's.
    CC 7.5
  3. 03
    CUDA error checkingEvery runtime call returns a status and every kernel launch fails silently unless you ask, which is what `cudaGetLastError` and a check macro are for.
    CC 7.5
  4. 04
    CUDA runtime APIThe cuda* calls in cudart that almost all CUDA code uses, layered on top of the driver API.
    CC 7.5
  5. 05
    cuda-gdbThe CUDA-aware debugger that can stop inside a kernel, select a GPU thread and inspect its device-side state.
    CC 7.5
  6. 06
    CUPTIThe profiling interface underneath Nsight and the PyTorch profiler, which is where the counter permission requirement comes from.
    CC 7.5
  7. 07
    lazy module loadingThe driver deferring each kernel's load until its first launch, which moves a cost most benchmarks then measure by accident.
    CC 7.5
  8. 08
    Nsight ComputeThe kernel profiler, which replays one kernel and reports its memory chart, its stall reasons and how close it got to the hardware limits.
    CC 7.5
  9. 09
    Nsight SystemsThe timeline profiler, which shows what the CPU, the copies and the kernels were each doing and when.
    CC 7.5
  10. 10
    nvccThe CUDA compiler driver, which splits your file into host code for the system compiler and device code for ptxas.
    CC 7.5
  11. 11
    NVTXAnnotations you add to your own code so a profiler timeline shows your phase names instead of anonymous bars.
    CC 7.5
  12. 12
    PTXNVIDIA's virtual instruction set, which ships inside your binary and gets compiled to real machine code for the GPU that runs it.
    CC 7.5
  13. 13
    ptxasThe assembler that turns PTX into SASS, decides how many registers each kernel gets, and reports when a kernel spills.
    CC 7.5
  14. 14
    PyTorch profilerPyTorch's operator-level profiler, which records CPU and CUDA activity, input shapes and traces that connect framework operations to device kernels.
    CC any CUDA GPU supported by the installed PyTorch build
  15. 15
    SASSThe actual machine code a GPU executes, which you read with cuobjdump or nvdisasm and which is where the truth lives.
    CC 7.5
05

12 terms

Hardware

12
  1. 01
    compute capabilityThe version number that says which features a GPU has, written 7.5 in docs and sm_75 on the nvcc command line.
    CC 7.5
  2. 02
    CUDA coreA lane of the SM's arithmetic pipeline, not a core in the CPU sense, and the number on the box is a count of these.
    CC any
  3. 03
    GPU architecture generationsTuring, Ampere, Ada, Hopper and Blackwell, each mapping to a range of compute capabilities and a set of features.
    CC 7.5
  4. 04
    green contextA CUDA execution context provisioned with a selected group of streaming multiprocessors, restricting its work to that SM partition.
    CC the measured binary targets 7.5; API availability depends on toolkit and driver
  5. 05
    host and deviceHost is the CPU and its memory, device is the GPU and its memory, and nothing crosses between them without a copy or a managed pointer.
    CC 7.5
  6. 06
    MPS and MIGTwo GPU-sharing mechanisms: MPS coordinates CUDA processes through a server, while MIG partitions supported datacenter GPUs into isolated device instances.
    CC MIG begins with supported datacenter Ampere GPUs; MPS feature floors vary
  7. 07
    PCIeThe bus between host memory and the GPU, roughly an order of magnitude slower than the GPU's own memory.
    CC 7.5
  8. 08
    register fileThe fixed pool of registers on each SM, partitioned among every thread resident there, which is what caps occupancy in most real kernels.
    CC 7.5
  9. 09
    resident warps and blocks per SMThe hard caps on how many warps and blocks one SM can hold at once, which set the ceiling occupancy can reach.
    CC 7.5
  10. 10
    streaming multiprocessorThe unit a GPU is actually made of: it owns a slice of registers and shared memory, holds many blocks at once, and issues warps from whichever one is ready.
    CC any
  11. 11
    tensor coreA specialized arithmetic unit that performs a small matrix multiply-accumulate cooperatively for a warp, at throughput ordinary scalar CUDA-core instructions cannot match.
    CC 7.0; lesson path requires 7.5
  12. 12
    warp schedulerThe part of the SM that picks, every cycle, which resident warp gets to issue an instruction.
    CC 7.5
06

15 terms

Precision

15
  1. 01
    2:4 structured sparsityKeeping exactly two values in every consecutive group of four so supported sparse tensor cores can skip a predictable half of a matrix.
    CC 8.0 for 2:4 sparse tensor-core acceleration
  2. 02
    accumulator typeThe data type used for a running sum or matrix output fragment, chosen independently from input storage to control rounding, range and hardware behavior.
    CC 7.5
  3. 03
    BF16A 16-bit float with FP32's eight exponent bits and seven stored mantissa bits, preserving range while giving up precision near ordinary values.
    CC 8.0 for tensor-core mma
  4. 04
    fast mathThe -use_fast_math flag, which swaps precise math functions for faster approximations and turns denormal support off.
    CC 7.5
  5. 05
    floating point determinismRepeatable floating-point results require a fixed input, reduction shape, addition order and compiled instruction sequence.
    CC 7.5
  6. 06
    FMAFused multiply-add: one instruction that computes a*b+c with a single rounding, which is both faster and more accurate than doing it in two steps.
    CC 7.5
  7. 07
    FP16A 16-bit floating-point format with five exponent bits and ten stored mantissa bits, giving compact storage but a maximum finite value of 65,504.
    CC 7.5
  8. 08
    FP32 and FP64The 32-bit single-precision and 64-bit double-precision IEEE formats, used respectively for ordinary GPU arithmetic and higher-accuracy references or scientific work.
    CC any
  9. 09
    FP8A family of eight-bit floating-point formats that trade range against precision, commonly using E4M3 for activations and E5M2 where wider range matters.
    CC 8.9 for tensor-core mma
  10. 10
    INT8 quantizationRepresenting real values as signed eight-bit integers plus scale and optional zero point, then accumulating integer products and rescaling the output.
    CC 7.5 for INT8 tensor cores; measured kernel uses CUDA cores
  11. 11
    mixed precisionComputing with narrow input or storage formats while retaining a wider accumulator, so bandwidth and arithmetic throughput improve without paying narrow-format error at every sum.
    CC 7.5
  12. 12
    mma.syncA PTX warp-level matrix instruction whose shape, layouts and data types determine exactly which registers each lane supplies and receives.
    CC 7.5 for m16n8k8 FP16
  13. 13
    MX block scalingAssigning one scale to a small block of low-precision values so each block can use the element format's limited range more closely.
    CC format-dependent; the day 98 T4 measurement is a host storage census
  14. 14
    TF32A tensor-core input format with FP32's range and at least ten bits of precision, used by Ampere-or-newer matrix instructions rather than ordinary CUDA-core arithmetic.
    CC 8.0
  15. 15
    WMMACUDA's warp-level matrix API, where 32 threads cooperatively load opaque fragments, execute a matrix multiply-accumulate, and store the resulting tile.
    CC 7.0; lesson path uses sm_75
07

15 terms

Libraries

15
  1. 01
    CCCLThe CUDA C++ Core Libraries, one repo holding Thrust, CUB and libcu++, versioned and shipped together.
    CC 7.5
  2. 02
    CUBThrust's engine, exposed at block, warp and device scope so you can drop a tuned primitive inside your own kernel.
    CC 7.5
  3. 03
    cuBLASNVIDIA's dense linear algebra library, providing architecture-tuned matrix operations whose data type, compute type and math mode define the numerical contract.
    CC any supported CUDA GPU
  4. 04
    cuBLASLtA descriptor-based matrix multiplication interface that exposes algorithm heuristics, workspace choices, layouts and fused epilogues beyond traditional cuBLAS GEMM calls.
    CC any supported CUDA GPU
  5. 05
    CUDA math librariesToolkit libraries that provide tuned random generation, transforms, sparse algebra and dense solvers, replacing custom kernels while exposing plans, state and scratch-buffer costs.
    CC any supported CUDA GPU
  6. 06
    cuDNNNVIDIA's deep-learning primitive library, whose graph API turns described operations into a selected engine, execution plan, workspace and bound tensor addresses.
    CC depends on operation and engine; measured FP32 graph ran at 7.5
  7. 07
    cuSPARSENVIDIA's sparse linear algebra library, providing descriptor-based operations such as CSR sparse matrix-vector multiply with algorithm and temporary-storage choices.
    CC any supported CUDA GPU
  8. 08
    CuTeCUTLASS's layout algebra and C++ type system, where a shape and stride define a coordinate-to-index function that can be composed and divided into tiled mappings.
    CC layout algebra is host-capable; measured device gather ran at 7.5
  9. 09
    CUTLASSNVIDIA's header-only CUDA C++ template library for composing architecture-specific matrix multiplication and related kernels from explicit data, tile, pipeline and epilogue choices.
    CC depends on the selected kernel; measured sm_75 path runs at 7.5
  10. 10
    gencodeThe nvcc target flags that choose which real-GPU machine code and virtual-architecture PTX images a CUDA binary carries.
    CC 7.5
  11. 11
    JIT compilationThe driver compiling embedded PTX into GPU machine code when a binary has no compatible prebuilt SASS image.
    CC 7.5
  12. 12
    libcu++The CUDA standard library: cuda::std types that work on both sides, plus device primitives like cuda::atomic_ref, cuda::barrier and cuda::pipeline.
    CC 7.5
  13. 13
    PyTorch custom operatorA user-defined operation registered with PyTorch's dispatcher, with explicit device implementations and optional fake and autograd registrations.
    CC set by the registered kernel; day 87 requires 7.5
  14. 14
    separate compilationCompiling device code in multiple translation units and joining it with a device-link step before the final host link.
    CC 7.5
  15. 15
    ThrustAn STL-shaped library of parallel algorithms over device data, where thrust::reduce replaces a kernel you would otherwise write.
    CC 7.5
08

2 terms

Ecosystem

2
  1. 01
    Compiler ExplorerThe site that compiles and runs CUDA in a browser on a real GPU, which is how every runnable snippet on this site works.
    CC 7.5
  2. 02
    free GPU tiersColab, Kaggle and Lightning give a GPU at no cost, with different cards, quotas and traps.
    CC 7.5