← Glossary
CUDA glossaryHardware
CC any

What is a CUDA core?

A lane of the SM's arithmetic pipeline, not a core in the CPU sense, and the number on the box is a count of these.

"Each CUDA core may only have one warp(32threads)", someone writes on r/CUDA, and the same idea turns up in the 86,520-view Stack Overflow question about how blocks, warps and threads map onto cores. It is the natural reading of the word and it is wrong in a way that costs you. A CUDA core has no instruction fetch of its own, no program counter, no cache attached and no ability to run anything by itself. It is one lane of arithmetic inside a streaming multiprocessor. What decides what happens next is the warp scheduler, and what it issues is a warp: one instruction for 32 threads.

So the number on the box is a count of arithmetic lanes across the whole chip, and putting it next to a CPU's core count compares two different things. A CPU core fetches, decodes, predicts and retires its own instruction stream; ten thousand CUDA cores share a few thousand of those front ends between them. Do not turn the count into a clock-by-clock model either: how many lanes an architecture lights up per cycle for a given instruction, and how many cycles a warp's instruction therefore takes, is an architecture detail that differs between FP32, FP64, integer and tensor core pipelines. cores x clock x 2 is a ceiling nobody's kernel reaches, and day 49 measures the gap on a real card.

The practical consequence is that you never address a lane. You fill a warp, and the SIMT width is 32 whether or not you asked for 32 threads. A 16-thread block still occupies a whole warp slot and still ties up 32 lanes, half of them doing nothing for the block's entire life. That is why two launches with the same thread count are not the same launch, and it is measurable on the smallest program in this course.

Measured

On a Tesla T4 (driver 595.84, CUDA 12.6, built with nvcc -O3 -arch=sm_75), day 2 launched 128 threads two ways and counted the lanes. As 8 blocks of 16 threads it takes 8 warp slots and leaves 128 of 256 lanes idle. As 4 blocks of 32 it takes 4 warp slots and leaves 0 of 128 idle. Same thread count, half the warp slots, and no scheduler anywhere merges two half-empty blocks into one warp. Captured 2026-08-30 on the project's verification node.

Diagram: what the marketing number counts. Left panel: one SM drawn as a warp scheduler box on top, feeding a row of 32 identical lane boxes below it, each lane containing only an arithmetic symbol and a register slice, with the labels "no instruction fetch", "no program counter", "no cache" running down the side. Right panel: the same 32 lanes twice. Top row, a 32-thread block, all lanes lit. Bottom row, a 16-thread block, 16 lanes lit and 16 greyed, with the caption "one warp slot either way". Alt text: "A CUDA core is one arithmetic lane under a shared warp scheduler, with no instruction fetch, program counter or cache of its own. A 16-thread block occupies a full 32-lane warp slot and leaves half its lanes idle."

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.