Day 2Module 1
in-technical-review

How a GPU differs from a CPU

This question turns up on r/CUDA in some form every few weeks, from people who have already written working kernels:

"each CUDA core may only have one warp(32threads) and each block is assigned to a SM (I am still confuse about what is it, if someone could explain what a SM is) ... what's the difference between me request 8blocks with 16thread, and 4 blocks with 32threads"

https://www.reddit.com/r/CUDA/comments/x2f767/how_does_cuda_blockswarps_thread_works/ (checked 2026-08-29)

Each noun there is a valid CUDA term, but the links between them are wrong. A CUDA core does not hold a warp, and a block is not a warp.

The two launches at the end also differ: both create 128 threads, but one uses twice as many warp slots.

This page answers it, lets you say whether one block can be split across two SMs, and explain why the same chip that loses on a small array wins on a big one.

What one SM actually contains

A GPU contains streaming multiprocessors, or SMs. Each SM has a register file, shared memory, warp schedulers, and arithmetic pipelines.

Those pipelines contain CUDA cores. A CUDA core is an arithmetic lane with no instruction fetch, program counter, or cache of its own. It is not comparable to a CPU core.

The thing the hardware schedules is a warp: "Each SM creates, manages, schedules, and executes threads in groups of 32 parallel threads called warps" (https://docs.nvidia.com/cuda/cuda-programming-guide/03-advanced/advanced-kernel-programming.html section 3.2.2.1, checked 2026-08-29). One instruction goes to all 32 lanes at once. That is SIMT, and it is why 32 turns up everywhere in this course.

The block you name in the launch configuration is not the scheduling unit, it is the allocation unit. When a block lands on an SM, "it partitions them into warps and each warp gets scheduled for execution by a warp scheduler", always the same way and always ceil(T / 32) of them (same page, section 3.2.2.2, checked 2026-08-29).

The hardware allocates whole warp slots. A block of 16 threads takes one slot and leaves 16 lanes idle. A block of 48 takes two slots and leaves 16 lanes idle in the second.

An SM has a fixed number of slots for each architecture. A Tesla T4 (compute capability 7.5) holds 32 resident warps and 16 resident blocks; an H100 holds 64 and 32.

Registers and shared memory are also "partitioned among the warps" and "partitioned among the thread blocks". The first of those four limits to bind sets your resident warp count. Tables 30 and 31 list the values: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/compute-capabilities.html (checked 2026-08-29).

The widget models one SM. Open its "what this does not show" list before trusting a number off it: no latency hiding and so no run time, no SM count or tail effect, no scheduler behaviour, no real register allocation, no spills, no clusters, no achieved occupancy.

A CPU avoids waiting, a GPU plans on it

A CPU spends most of its transistors keeping one instruction stream from stalling: deep caches, out-of-order execution, branch prediction. That is expensive per thread, so it keeps few threads in flight and pays a real cost to swap between them, because registers and pipeline state have to be saved and restored.

A GPU keeps each warp's execution state on the SM.

The guide says: "The execution context (program counters, registers, etc.) for each warp processed by an SM is maintained on-chip throughout the warp's lifetime. Therefore, switching between warps incurs no cost" (same page, section 3.2.2.2, checked 2026-08-29).

The scheduler does not need to save or restore that state.

When one warp waits on global memory, the warp scheduler can issue an instruction from another resident warp. This does not remove the stall; it keeps another warp running during it.

CUDA calls this latency hiding. Occupancy measures the active warps or threads as a share of the SM's limit.

The memory hierarchy matches these scopes. Registers belong to a thread, shared memory belongs to a block on an SM, L1 belongs to an SM and L2 serves the device, and global memory is off-chip DRAM.

A CPU relies on caches to avoid DRAM access. A GPU often runs another warp while one waits for DRAM, and CUDA lets you load shared memory yourself. Days 11 to 20 cover these memory spaces.

So on a small problem the GPU loses, and that is the design working rather than a bug in your code. Day 9 finds the crossover.

Three CPU habits that break here

A thread is not independent. It runs as one lane in a 32-thread warp that shares an instruction stream. If a branch splits the warp, the hardware runs each path with some lanes masked; day 22 measures this warp divergence.

Do not assume that a warp always moves in lockstep. Since compute capability 7.0, "independent thread scheduling allows full concurrency between threads, regardless of warp", and "the ability for threads to diverge and reconverge at sub-warp granularity makes such assumptions invalid" (same page, section 3.2.2.1.1, checked 2026-08-29).

Independent thread scheduling is why warp primitives have a _sync suffix and a mask argument.

The scheduler cannot move part of a block to another SM. An operating system can move CPU threads between cores, but CUDA assigns each block to one SM. The guide states that "all threads of a block reside on the same streaming multiprocessor(SM) and must share the resources of the SM" (https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/intro-to-cuda-cpp.html section 2.1.2.2, checked 2026-08-29).

This scope is why __syncthreads() is block-wide. Day 28 introduces grid-wide sync.

More threads is not automatically better. The grid is a queue, not a promise: what does not fit waits.

Getting these numbers off your own card

Full program in code/day02-gpu-vs-cpu/warp_slots.cu. Everything above is a claim about hardware state, and the point of the program is that you stop taking it on my word.

Three rules it follows.

Every per-SM limit comes from the driver. cudaGetDeviceProperties reports warp size, max threads per block, resident threads and blocks per SM, registers per SM, shared memory, L2 size. No table copied off a blog, including this one.

The residency figure is the driver's answer, not our arithmetic. cudaOccupancyMaxActiveBlocksPerMultiprocessor takes the compiled kernel and a block size and returns how many blocks fit (https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__OCCUPANCY.html , checked 2026-08-29). It accounts for the register count, which your source never shows you and which changes when you edit the kernel body.

Nothing is launched and nothing is timed, so no missing warm-up or host clock can spoil a number. Timing starts on day 9.

The warp count is the guide's formula and nothing else:

// Warp slots a block of t threads occupies, which is ceil(t / 32). This is the
// guide's own formula, and it rounds up: a block of 48 threads takes two slots
// and the second one runs with 16 of its 32 lanes switched off.
// constexpr so the two claims below can be checked at compile time. They
// follow only from the constants in this file, so a run-time check would look
// like a measurement of your GPU and be nothing of the kind.
constexpr int warpsPerBlock(int t) {
    return (t + kWarpSize - 1) / kWarpSize;
}

And the loop that asks the driver, with the two caps this page claims checked against the driver rather than promised:

    for (int c = 0; c < kNumBlockSizes; ++c) {
        const int t = kBlockSizes[c];
        const int warps = warpsPerBlock(t);
        const int idle = warps * kWarpSize - t;

        int resident = 0;
        CUDA_CHECK(cudaOccupancyMaxActiveBlocksPerMultiprocessor(
            &resident, scaleAdd, t, 0));
        const int residentWarps = resident * warps;

        // The two caps the page states in words, checked against the driver's
        // own answer rather than asserted at the reader. A residency above
        // either of them would mean the model on the page is wrong.
        //
        // These are real branches, not asserts. CI builds Release, which
        // defines NDEBUG, and an assert under NDEBUG is deleted. A check that
        // vanishes in the build that matters would let a wrong table reach the
        // page with CI green, which is the one outcome this program exists to
        // prevent.
        if (resident > prop.maxBlocksPerMultiProcessor) {
            std::fprintf(stderr,
                         "residency %d exceeds the per-SM block cap %d at %d "
                         "threads per block\n",
                         resident, prop.maxBlocksPerMultiProcessor, t);
            ++wrong;
        }
        if (residentWarps > maxWarpsPerSM) {
            std::fprintf(stderr,
                         "resident warps %d exceeds the per-SM warp cap %d at "
                         "%d threads per block\n",
                         residentWarps, maxWarpsPerSM, t);
            ++wrong;
        }

        std::printf("%11d %10d %11d %11d %10d %9.0f%%\n", t, warps, idle,
                    resident, residentWarps,
                    100.0 * residentWarps / maxWarpsPerSM);
    }

The occupancy API reports values for scaleAdd, a one-multiply-one-add kernel that uses few registers. Your kernel may have lower occupancy because it uses more registers.

Day 17 reads the register count from ptxas. Day 45 shows why maximum occupancy is not always the best target.

Results

Re-verified without behavioral drift on a Tesla T4 with driver 580.173.02 and CUDA 13.0 (V13.0.88) on 2026-09-02. The original CUDA 12.6 transcript remains beside the new run in the page's evidence array.

One SM, one block size at a time. The blocks/SM column is what the
driver reports for scaleAdd, not a formula this program made up.

threads/blk  warps/blk  idle lanes   blocks/SM   warps/SM  occupancy
         16          1          16          16         16        50%
         32          1           0          16         16        50%
         48          2          16          16         32       100%
         64          2           0          16         32       100%
        128          4           0           8         32       100%
        256          8           0           4         32       100%
        512         16           0           2         32       100%
       1024         32           0           1         32       100%

128 threads, launched two ways
  8 blocks x 16 threads: 8 warp slots, 128 of 256 lanes idle
  4 blocks x 32 threads: 4 warp slots, 0 of 128 lanes idle

Every row above describes one SM. Multiply blocks/SM by the SM
count to see how much of a grid is resident at once; the rest of the
grid is queued, not running.

The prediction below was made before the run, and it held. It said that the block limit binds first at 16 and 32 threads, so 16 blocks fill only 16 of 32 warp slots. From 33 to 64 threads, the block and warp limits meet.

The measured table shows 50 percent occupancy at 16 and 32 threads, then 100 percent from 48 threads up. blocks/SM stays at 16 through 64 threads and then falls.

Tables 30 and 31 supplied the limits used in the prediction. On the tested GPU, each SM has 32 warp slots and 16 block slots. A 16-thread or 32-thread block takes one warp slot, so the block limit leaves half the warp slots empty.

From 33 to 64 threads, the two limits meet. Above that, the warp limit binds first. An architecture with 64 warp slots and 32 block slots reaches the same crossing at a different point.

Do not memorise a residency, look for which cap binds. That is the number that tells you what to change.

Run it yourself

You do not need a GPU today. The quiz below is answerable from this page, which is why this day's minimum compute capability is any.

If you have one, the program builds and runs anywhere CUDA 13 does:

nvcc -std=c++17 -O3 -arch=sm_75 -o warp_slots warp_slots.cu
./warp_slots

There is no Compiler Explorer embed on this page because the program exceeds the 60-line embed limit. You can paste it into https://godbolt.org/ (checked 2026-08-29) and target sm_75 or lower.

Colab also works. If you have neither, read /setup/learn-cuda-without-a-gpu and use the widget.

Exercise

Three questions, all answerable from the sections above.

  1. You launch 8 blocks of 16 threads, someone else launches 4 blocks of 32. Both are 128 threads. Which uses the GPU better, and why?
  2. Your kernel wants 96 KB of shared memory per block and your card offers 64 KB per SM. Can one block split across two SMs so it fits?
  3. Your card has 40 SMs and holds 16 blocks per SM. You launch 5,000 blocks of 256 threads. How many run the instant the kernel starts?

Time: 15 minutes. Submit: nothing.

Check: the questions and their option notes live in content/quizzes/day02.toml, and the browser marks your choices. A wrong answer opens the note for every option. The site does not record a score.

If you ran warp_slots.cu, its last output block answers question 1 from your card.

Hint 1

Count what the hardware allocates, not what you asked for. You asked for threads. What does the SM hand out?

Hint 2

ceil(T / 32), and it rounds up. Apply it to a 16-thread block and then to a 32-thread block, and multiply each by the block count.

Solution

1. 4 blocks of 32. The hardware allocates whole warp slots, so a 16-thread block takes one slot and leaves 16 lanes idle.

Eight such blocks use eight slots for 128 threads; four 32-thread blocks use four slots.

The answer "more blocks means more parallelism" treats blocks as the scheduling unit. Warps are the scheduling unit, and the scheduler does not merge two half-full blocks into one warp.

2. No, and it will not launch. All threads in a block reside on one SM.

The guide states: "If there are not enough resources available per SM to process at least one block, the kernel will fail to launch" (https://docs.nvidia.com/cuda/cuda-programming-guide/03-advanced/advanced-kernel-programming.html section 3.2.2.2, checked 2026-08-29).

A block cannot use the combined shared memory of two SMs.

3. At most 640, and in fact fewer. 40 SMs times 16 blocks is 640, so the other 4,360 wait.

But 16 blocks of 256 threads need 128 warp slots on an SM that has 32 to 64, depending on the architecture (see FACT-SHEET.md section 3).

The warp limit therefore binds first. The grid contains all 5,000 blocks, but only the resident blocks run at once.

A launch configuration is a request for warp slots, not for threads.

Pitfalls

Reading the CUDA core count as a CPU core count. A CUDA core is an arithmetic lane with no instruction fetch or program counter. Comparing ten thousand CUDA cores with 16 CPU cores does not compare equivalent parts.

Day 49 compares bandwidth and resident warps instead.

Computing occupancy as blocks times threads divided by 32. That formula appears in NVIDIA's occupancy post (https://developer.nvidia.com/blog/cuda-pro-tip-occupancy-api-simplifies-launch-configuration/ , checked 2026-08-29). It works when the block size is a multiple of 32.

At 16 threads per block, it counts half a warp even though the SM allocates a whole slot. Use blocks times ceil(T / 32).

Asking for more threads per block than the device allows. The product of the block dimensions must stay at or below maxThreadsPerBlock, which is 1024 on every compute capability this course covers. A dim3(32, 32, 2) block asks for 2048 threads and returns invalid configuration argument.

Day 4 covers the arithmetic, and the error page lists other causes.

Treating occupancy as the score to maximise. Occupancy gives the scheduler warps to run during a stall. Once there are enough warps, adding more may not help and can reduce the registers available to each thread.

Day 45 measures two configurations that run at the same speed with different occupancy.

Go deeper

Next

Day 3 gets CUDA working on the machine you actually own, or hands you a free one. Day 4 turns the block and grid you just learned to count into the index arithmetic that decides which thread touches which element, which is where the next batch of wrong answers lives.