What is a warp shuffle in CUDA?
Instructions that let lanes of one warp read each other's registers directly, with no shared memory and no barrier.
A shuffle is a read with no address. __shfl_down_sync(mask, value, 16) hands lane 0 whatever lane 16 holds in its own copy of value, and nothing is stored, loaded or allocated on the way: the guide's sentence is that these functions "exchange a value between non-exited threads within a warp without the use of shared memory". There are four, and they differ only in how each lane names its source: __shfl_sync a lane you name, which is a broadcast when all 32 name the same one, __shfl_down_sync and __shfl_up_sync lane L + delta and L - delta with no wrapping, __shfl_xor_sync lane L xor laneMask.
Which one you pick decides where the answer lands, not how fast you get it. Five __shfl_down_sync steps at deltas 16, 8, 4, 2 and 1 fold a warp's 32 values into lane 0 and leave the other 31 holding partial sums nobody wants, so a kernel whose lanes all need the total pays a sixth instruction to broadcast it back. Five __shfl_xor_sync steps cost the same five and finish with the answer in all 32.
The mask is where the bugs are, and there are two rules where most people learn one. The first is membership: the mask names the lanes that must reach this instruction, the hardware waits for them, each calling lane sets its bit, each non-calling lane clears it, and everybody passes the same value. Break that and the guide says the behaviour is "invalid, such as kernel hang, or undefined". The second rule is about the lane you read rather than the lanes you named: "Threads may only read data from another thread that is actively participating in the intrinsics. If the target thread is inactive, the retrieved value is undefined." A mask can satisfy the first rule and the read still be undefined. Day 23 ships that kernel, because it is the canonical example: __ballot_sync over 20 live columns returns 0x000fffff, which is the mask every tutorial tells you to compute, and at the first step lane 4 asks for lane 20, which that mask does not name.
So the fix for a partly live warp is usually not a narrower mask. Put every lane back inside the collective and give the ones with no data the identity of the operation: 0 for a sum, a large negative float for a max, false for __any_sync, true for __all_sync. Now 0xffffffff is honest, every source lane is live, and the answer is defined. Guard the loads and the stores, never the collective, which is __syncthreads()'s rule from day 14 turning up where there is no barrier at all. Since compute capability 7.0 each lane carries its own program counter, and independent thread scheduling is why these take a mask in the first place.
Measured
Day 23 gives one warp one row: 8,205 rows of 20 live columns in a 32-column stride, 256 threads per block, 1,026 blocks, 8,208 warps launched of which 3 own no row. Then it asks the driver what each version costs. Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75.
| Kernel | Shared per block | Registers per thread | Rows wrong of 8,205 |
|---|---|---|---|
reduceRowsShared |
2048 B | 15 | 0 |
reduceRowsPadded, shuffles |
0 B | 23 | 0 |
reduceRowsBallotDown |
0 B | 16 | 0 |
Those come from cudaFuncGetAttributes, so the claim is the driver's rather than the page's. The shuffle kernels reserve no shared memory and still get every row right, and shared memory is the resource that decides how many blocks fit on an SM. It is not free: they want 23 and 16 registers per thread where the shared version wants 15, shared memory traded for register pressure.
Whether it buys time is another question, and here the answer is no. Day 25 stops a block reduction's shared tree at 64 live values and finishes with __shfl_down_sync, over 67,108,864 floats: 1.132 ms and 237.1 GB/s against 1.125 ms and 238.5 GB/s for the cascaded version it was built from. That kernel is memory bound from the cascade onward, so deleting instructions from its last six steps changes nothing a clock can see. Both figures count only the 256 MiB each kernel reads, which is why they sit above the copy kernel's 201.2 GB/s.
Diagram
An original SVG, warp-shuffle-ladder-vs-butterfly. Two panels of 32 lane boxes with arrows and no memory drawn anywhere. The left panel is three bands of __shfl_down_sync at offsets 16, 4 and 1, arrows running right to left, the number of lanes still holding a partial sum falling 16, 4, 1. The right panel is the same three bands of __shfl_xor_sync, arrows crossing in both directions, every lane holding the running answer at every step.
Alt text: "Two ways a warp reduces itself in five steps. The down-shuffle ladder leaves the total in lane zero alone; the xor butterfly costs the same five steps and leaves it in all thirty-two lanes. Neither picture contains any memory."
Code
From code/day23-shuffles/shuffles.cu. This is the whole warp sum, and the second half is the part a diagram cannot make obvious: after the ladder, 31 of the 32 lanes are holding nothing you want.
float sum = value;
for (unsigned int offset = kWarpSize / 2; offset > 0; offset /= 2) {
sum += __shfl_down_sync(kFullMask, sum, offset);
}
// Every lane needs the total to normalise its own value, and only lane 0
// has it, so one broadcast copies lane 0's register into all 32.
const float rowSum = __shfl_sync(kFullMask, sum, 0);
kFullMask is 0xffffffff and it is correct here for one reason: the lanes with no row data reached the same instruction carrying 0, the identity of the sum.
Related terms
Where you meet this
- Day 21, what a warp is, which reads the lane numbering out of the hardware that these instructions address.
- Day 22, warp divergence, where a split warp costs throughput, one page before it starts costing correctness.
- Day 23,
__shfl_down_syncand the mask argument, the lesson that owns this term. - Day 25, parallel reduction part 2, where the shuffle tail replaces the
volatileone and the clock does not move. - Colab setup, one of the free T4s day 23 runs on, alongside Compiler Explorer, which executes it.
Sources
- CUDA Programming Guide 5.4.6 "Warp Functions" and 5.4.6.6 for the mask constraints and the source-lane rule quoted above: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/cpp-language-extensions.html (checked 2026-08-30)
- "Using CUDA Warp-Level Primitives", NVIDIA's own post, including the canonical reduction whose mask is legal and whose reads are not: https://developer.nvidia.com/blog/using-cuda-warp-level-primitives/ (checked 2026-08-30)
- CUDA Programming Guide 3.2.2.1.1, independent thread scheduling, for why every one of these takes a mask: https://docs.nvidia.com/cuda/cuda-programming-guide/03-advanced/advanced-kernel-programming.html (checked 2026-08-30)
cuda-samples,shfl_scan, NVIDIA's shuffle sample: https://github.com/NVIDIA/cuda-samples/tree/master/cpp/2_Concepts_and_Techniques/shfl_scan (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.