← Glossary
CUDA glossaryExecution model
CC 7.0

What is independent thread scheduling in CUDA?

Since Volta, each thread has its own program counter, so lanes of a warp can be at different instructions and the old lockstep assumptions break.

The _sync on the end of __shfl_down_sync is the visible edge of a hardware change. Below compute capability 7.0 a warp had one program counter, "shared amongst all 32 threads in the warp together with an active mask specifying the active threads of the warp", and the mask was the only per-thread state in the picture. From 7.0 the guide describes something else: "the GPU maintains execution state per thread, including a program counter and call stack, and can yield execution at a per-thread granularity". A schedule optimizer still gathers the active threads back into SIMT groups, so the throughput survives. What did not survive is the guarantee.

What that costs you is every inference from one lane's progress to another's. The guide names the casualty: "Warp-synchronous code assumes that threads in the same warp execute in lockstep at every instruction, but the ability for threads to diverge and reconverge at sub-warp granularity makes such assumptions invalid." Reconvergence at the closing brace of an if is no longer implied either, which is why __syncwarp() exists as the warp-scoped sibling of __syncthreads(). And the participants in a collective became yours to name rather than the hardware's to assume: "The set of threads that participates in invoking each primitive is specified using a 32-bit mask, which is the first argument of these primitives."

Naming them correctly is where careful people still lose. The mask rules are short: every calling lane sets its own bit, and every lane named in the mask has to reach that same instruction. A shuffle carries a second rule under those, that a thread "may only read data from another thread that is actively participating in the intrinsics", and no mask enforces it for you. Build the mask with __ballot_sync over a live predicate, exactly as the canonical example does, put a __shfl_down_sync ladder under it, and at offset 16 lane 4 asks lane 20 for a value. When the row is 20 columns wide, lane 20 is not in the mask and is not running that line. Every mask rule holds. The read is undefined.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75. Captured 2026-08-30 on the project's verification node.

Day 23 ships that broken ladder on purpose, over 8205 rows of 20 live columns so the ballot returns 0x000fffff and never a power of two. It got 0 sums wrong. The evidence file refuses to treat that as a defence, in its own words: "every row came out right, which is a fact about this driver and not about the code". Undefined is not the same as noisy. The answer arrives, it looks right, and it is not yours to depend on. Day 23 prints those numbers and gates none of them, because a test demanding the wrong value would be asserting that undefined behaviour is dependable.

Day 21 measures the other half on the same card. Every one of a warp's 32 lanes reported the same clock64() value, and no two warps reported the same one, so lockstep is exactly what this T4 did. A T4 is compute capability 7.5, so per-thread program counters were on the entire time. That is the shape of the trap: the hardware keeps behaving like the old model right up to the build where it stops. Day 22, which owns this term, prices the diverged state itself at 3.272 ms against 1.616 ms for the same arithmetic split on a warp boundary.

Code

From code/day23-shuffles/shuffles.cu. The legal mask and the undefined read, five lines apart.

// Every lane of the warp reaches the ballot, so kFullMask is right here.
// liveMask comes back as 0x000fffff, lanes 0 to 19, which is a legal mask
// and is what the canonical example tells you to pass.
const unsigned int liveMask = __ballot_sync(kFullMask, live);
if (live) {
    float sum = in[row * kStride + lane];
    for (unsigned int offset = kWarpSize / 2; offset > 0; offset /= 2) {
        sum += __shfl_down_sync(liveMask, sum, offset);
    }

The fix is not a wider mask. It is to keep all 32 lanes inside the collective and have the lanes with no data carry the identity of the operation, which is what the working kernel on the same page does.

Diagram

An original SVG: two warps drawn as 32 lane boxes on a time axis. Above, the pre-Volta warp, one program counter arrow serving all 32 boxes with a mask strip below it. Under it, the Volta-and-later warp, 32 separate arrows, four of them parked at a different instruction than the rest, and a __syncwarp() bar where they line back up.

Alt text: "Before Volta one program counter served all 32 lanes of a warp and a mask said which were active. From Volta each lane carries its own program counter, so lanes can sit at different instructions until a __syncwarp() brings them back together."

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.