← Glossary
CUDA glossaryExecution model
CC 7.5

How do I compute a global thread index in CUDA?

The single number that identifies a thread across the whole grid, usually blockIdx.x * blockDim.x + threadIdx.x.

Read the formula as a sentence and it stops being a line you copy. Count the blocks ahead of mine, multiply by how many threads each one holds, add my seat number inside my own block. It runs backwards just as easily, which is the version that proves you understand it: given an element i, its block is i / blockDim.x and its thread is i % blockDim.x. With 256 threads per block, element 1337 is block 5, thread 57, because 1337 = 5 * 256 + 57. With 1024, the same element is block 1, thread 313. Neither answer is a fact about CUDA. Both are facts about a launch.

The same shape covers more dimensions, and only the flattening changes:

Shape Index
1D blockIdx.x * blockDim.x + threadIdx.x
2D, row-major image of width w row * w + col, where col and row are each built the 1D way on .x and .y
3D, volume of width w and height h (z * h + y) * w + x, with x, y and z each built the 1D way

Note what the 2D row does not say. It does not say col * h + row. Which of the two you want is decided by how your array is laid out in memory, not by CUDA, and getting it backwards gives a transposed image that is still fully populated, so no bounds check and no error string will find it. Day 7 is where that choice gets made against a real buffer.

Three ways the 1D version goes wrong, in the order they cost you. Swapping blockDim for gridDim produces a legal index that is far too small: under <<<8, 256>>> the multiplier becomes 8 instead of 256, so every thread lands in the first 312 elements and the rest of the array keeps whatever it held. Nothing fails; the answer is just wrong. Dropping the if (i < n) guard is the opposite, an index past the end, and day 6 measured where it shows up: cudaSuccess at the launch, then an illegal memory access was encountered at the next cudaMemcpy, and after that every call on the context returns the same error, including a fresh cudaMalloc and both cudaFree calls. The error names the copy, not the kernel. Storing the index in the wrong type is the one that waits. blockIdx.x and blockDim.x are both unsigned 32-bit, so their product is computed in 32 bits no matter what you assign it to; a signed int result is already wrong above 2^31 and the unsigned product itself wraps at 2^32, which is 4 Gi elements and fits on cards people rent by the hour. Widening after the multiply changes nothing. Cast blockDim.x to size_t before it.

Measured

The run behind this: a Tesla T4 on driver 595.84, CUDA 12.6 (V12.6.85), nvcc -O3 -arch=sm_75, on the project's verification node, 2026-08-30.

Day 4 does not time anything. It checks the formula against the hardware, which is a different kind of measurement and the only one this term needs. Each thread writes its own blockIdx.x and threadIdx.x into the element it owns; the host then rebuilds block * blockDim + thread for every written element and fails on the first one that is not its own index. Over 2000 elements under <<<8, 256>>>, all 2000 agreed.

Launch Element 1337 resolves to Lane
<<<8, 256>>> block 5, thread 57 25
<<<2, 1024>>> block 1, thread 313 25

The lane holding still across both is worth a second look. A thread's lane is i % 32 whenever the block size is a whole number of warps, so it is a property of the element and not of the launch. That is the reason the global index is the right thing to keep consecutive across a warp: consecutive indices are consecutive addresses, which is what memory coalescing rewards, and day 11 measured a 232.9 GB/s copy fall to 9.6 GB/s when they were not.

The formula also stops being the whole story once one thread owns more than one element. Day 8 starts the index here and then steps it by blockDim.x * gridDim.x, which is the only place gridDim legitimately enters the arithmetic.

Diagram

index-tracer, preset 1d-basic, in reverse mode: click an element and the widget draws the arrow back to its owning thread and prints the division and the remainder that got there.

Alt text: "An element clicked in a twelve-cell array, with an arrow drawn back to block 1, thread 1, and the arithmetic 5 divided by 4 is 1 remainder 1 shown beside it."

Related terms

Where you meet this

Sources

Byline

Written by: unassigned. Reviewed by: unassigned. Written on: not set. Last checked: not set. Numbers captured 2026-08-30 on the project's verification node. This entry stays a draft until a named author and a different named reviewer sign it.