What is independent thread scheduling in CUDA?
Since Volta, each thread has its own program counter, so lanes of a warp can be at different instructions and the old lockstep assumptions break.
The _sync on the end of __shfl_down_sync is the visible edge of a hardware change. Below compute capability 7.0 a warp had one program counter, "shared amongst all 32 threads in the warp together with an active mask specifying the active threads of the warp", and the mask was the only per-thread state in the picture. From 7.0 the guide describes something else: "the GPU maintains execution state per thread, including a program counter and call stack, and can yield execution at a per-thread granularity". A schedule optimizer still gathers the active threads back into SIMT groups, so the throughput survives. What did not survive is the guarantee.
What that costs you is every inference from one lane's progress to another's. The guide names the casualty: "Warp-synchronous code assumes that threads in the same warp execute in lockstep at every instruction, but the ability for threads to diverge and reconverge at sub-warp granularity makes such assumptions invalid." Reconvergence at the closing brace of an if is no longer implied either, which is why __syncwarp() exists as the warp-scoped sibling of __syncthreads(). And the participants in a collective became yours to name rather than the hardware's to assume: "The set of threads that participates in invoking each primitive is specified using a 32-bit mask, which is the first argument of these primitives."
Naming them correctly is where careful people still lose. The mask rules are short: every calling lane sets its own bit, and every lane named in the mask has to reach that same instruction. A shuffle carries a second rule under those, that a thread "may only read data from another thread that is actively participating in the intrinsics", and no mask enforces it for you. Build the mask with __ballot_sync over a live predicate, exactly as the canonical example does, put a __shfl_down_sync ladder under it, and at offset 16 lane 4 asks lane 20 for a value. When the row is 20 columns wide, lane 20 is not in the mask and is not running that line. Every mask rule holds. The read is undefined.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75. Captured 2026-08-30 on the project's verification node.
Day 23 ships that broken ladder on purpose, over 8205 rows of 20 live columns so the ballot returns 0x000fffff and never a power of two. It got 0 sums wrong. The evidence file refuses to treat that as a defence, in its own words: "every row came out right, which is a fact about this driver and not about the code". Undefined is not the same as noisy. The answer arrives, it looks right, and it is not yours to depend on. Day 23 prints those numbers and gates none of them, because a test demanding the wrong value would be asserting that undefined behaviour is dependable.
Day 21 measures the other half on the same card. Every one of a warp's 32 lanes reported the same clock64() value, and no two warps reported the same one, so lockstep is exactly what this T4 did. A T4 is compute capability 7.5, so per-thread program counters were on the entire time. That is the shape of the trap: the hardware keeps behaving like the old model right up to the build where it stops. Day 22, which owns this term, prices the diverged state itself at 3.272 ms against 1.616 ms for the same arithmetic split on a warp boundary.
Code
From code/day23-shuffles/shuffles.cu. The legal mask and the undefined read, five lines apart.
// Every lane of the warp reaches the ballot, so kFullMask is right here.
// liveMask comes back as 0x000fffff, lanes 0 to 19, which is a legal mask
// and is what the canonical example tells you to pass.
const unsigned int liveMask = __ballot_sync(kFullMask, live);
if (live) {
float sum = in[row * kStride + lane];
for (unsigned int offset = kWarpSize / 2; offset > 0; offset /= 2) {
sum += __shfl_down_sync(liveMask, sum, offset);
}
The fix is not a wider mask. It is to keep all 32 lanes inside the collective and have the lanes with no data carry the identity of the operation, which is what the working kernel on the same page does.
Diagram
An original SVG: two warps drawn as 32 lane boxes on a time axis. Above, the pre-Volta warp, one program counter arrow serving all 32 boxes with a mask strip below it. Under it, the Volta-and-later warp, 32 separate arrows, four of them parked at a different instruction than the rest, and a __syncwarp() bar where they line back up.
Alt text: "Before Volta one program counter served all 32 lanes of a warp and a mask said which were active. From Volta each lane carries its own program counter, so lanes can sit at different instructions until a __syncwarp() brings them back together."
Related terms
Where you meet this
- Day 21, what is a warp, where a T4 shows you lockstep on a card that does not promise it.
- Day 22, warp divergence, the lesson that owns this term.
- Day 23, warp shuffles, where you write your first mask and meet the ladder above.
- Day 24, parallel reduction, whose warp-synchronous tail is the classic piece of code this change invalidated.
- Colab setup, a compute capability 7.5 card, so everything on this page is on by default there.
Sources
- CUDA Programming Guide 3.2.2.1.1, "Independent Thread Scheduling", for the per-thread program counter, the pre-7.0 model and the warning on warp-synchronous code: https://docs.nvidia.com/cuda/cuda-programming-guide/03-advanced/advanced-kernel-programming.html (checked 2026-08-30)
- CUDA C++ Best Practices Guide 13.1, "Branching and Divergence", for a warp remaining diverged past the conditional block and for
__syncwarp(): https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html (checked 2026-08-30) - CUDA Programming Guide, C++ language extensions, for the warp shuffle functions and the rule that a thread may only read from a lane actively participating in the intrinsic: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/cpp-language-extensions.html (checked 2026-08-30)
- "Using CUDA Warp-Level Primitives", NVIDIA developer blog, for the mask as the first argument of every primitive and for the
__ballot_syncidiom the trap above is built from: https://developer.nvidia.com/blog/using-cuda-warp-level-primitives/ (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.