What is a CUDA core?
A lane of the SM's arithmetic pipeline, not a core in the CPU sense, and the number on the box is a count of these.
"Each CUDA core may only have one warp(32threads)", someone writes on r/CUDA, and the same idea turns up in the 86,520-view Stack Overflow question about how blocks, warps and threads map onto cores. It is the natural reading of the word and it is wrong in a way that costs you. A CUDA core has no instruction fetch of its own, no program counter, no cache attached and no ability to run anything by itself. It is one lane of arithmetic inside a streaming multiprocessor. What decides what happens next is the warp scheduler, and what it issues is a warp: one instruction for 32 threads.
So the number on the box is a count of arithmetic lanes across the whole chip, and putting it next to a CPU's core count compares two different things. A CPU core fetches, decodes, predicts and retires its own instruction stream; ten thousand CUDA cores share a few thousand of those front ends between them. Do not turn the count into a clock-by-clock model either: how many lanes an architecture lights up per cycle for a given instruction, and how many cycles a warp's instruction therefore takes, is an architecture detail that differs between FP32, FP64, integer and tensor core pipelines. cores x clock x 2 is a ceiling nobody's kernel reaches, and day 49 measures the gap on a real card.
The practical consequence is that you never address a lane. You fill a warp, and the SIMT width is 32 whether or not you asked for 32 threads. A 16-thread block still occupies a whole warp slot and still ties up 32 lanes, half of them doing nothing for the block's entire life. That is why two launches with the same thread count are not the same launch, and it is measurable on the smallest program in this course.
Measured
On a Tesla T4 (driver 595.84, CUDA 12.6, built with nvcc -O3 -arch=sm_75), day 2 launched 128 threads two ways and counted the lanes. As 8 blocks of 16 threads it takes 8 warp slots and leaves 128 of 256 lanes idle. As 4 blocks of 32 it takes 4 warp slots and leaves 0 of 128 idle. Same thread count, half the warp slots, and no scheduler anywhere merges two half-empty blocks into one warp. Captured 2026-08-30 on the project's verification node.
Diagram: what the marketing number counts. Left panel: one SM drawn as a warp scheduler box on top, feeding a row of 32 identical lane boxes below it, each lane containing only an arithmetic symbol and a register slice, with the labels "no instruction fetch", "no program counter", "no cache" running down the side. Right panel: the same 32 lanes twice. Top row, a 32-thread block, all lanes lit. Bottom row, a 16-thread block, 16 lanes lit and 16 greyed, with the caption "one warp slot either way". Alt text: "A CUDA core is one arithmetic lane under a shared warp scheduler, with no instruction fetch, program counter or cache of its own. A 16-thread block occupies a full 32-lane warp slot and leaves half its lanes idle."
Related terms
Where you meet this
- Day 2, how a GPU differs from a CPU, which owns the numbers above.
- Day 10, threads per block, where idle lanes turn into time on a real kernel.
- Day 49, the roofline, which measures how far a kernel gets from the peak the core count implies.
- Monitoring GPU usage, because
nvidia-smireports a utilization figure, not a count of busy lanes.
Sources
- CUDA Programming Guide, "SIMT Architecture", on the SM scheduling threads in groups of 32: https://docs.nvidia.com/cuda/cuda-programming-guide/03-advanced/advanced-kernel-programming.html (checked 2026-08-29)
- "How do CUDA blocks/warps/threads map onto CUDA cores?", 86,520 views: https://stackoverflow.com/questions/10460742/how-do-cuda-blocks-warps-threads-map-onto-cuda-cores (checked 2026-08-29)
- The r/CUDA thread quoted above: https://www.reddit.com/r/CUDA/comments/x2f767/how_does_cuda_blockswarps_thread_works/ (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.