← Glossary
CUDA glossaryMemory
CC 7.5

What is a bank conflict in CUDA shared memory?

Two lanes of a warp hitting different addresses in the same bank, which serializes their accesses.

You write one shared read, and the memory answers it in as many passes as its busiest bank needs. NVIDIA states the rule without hedging: "If multiple addresses of a memory request map to the same memory bank, the accesses are serialized. The hardware splits a memory request that has bank conflicts into as many separate conflict-free requests as necessary, decreasing the effective bandwidth by a factor equal to the number of separate memory requests." So the quantity worth computing is the conflict degree, the largest number of distinct words any single bank is asked for. Degree 1 is one request. Degree 32 is 32 requests for an instruction you issued once.

For the pattern that causes most of them there is a closed form. Lane L reading word L * s lands in bank L * s mod 32; two lanes collide when their index difference times s is a multiple of 32, which happens every 32 / gcd(s, 32) lanes, so each occupied bank collects gcd(s, 32) lanes and their words all differ. The degree is gcd(stride, 32), and the size of the stride has nothing to do with it. That is the opposite of what day 11 taught about global memory, where every step up the stride cost more bandwidth than the last. Shared memory is periodic rather than monotone: strides 31, 33, 63 and 65 are all free, and only the factors a stride shares with 32 are ever charged for.

The usual way to get this wrong is to time a broadcast and report it as a conflict. Lanes reading the same address are "coalesced into a single multicast", not serialized, and the distinction is where a widely copied demo falls over. A published 120-day CUDA challenge indexes its slow kernel with tx * 2 % 32, which puts lanes tx and tx + 16 on one address: the writes race, the reads broadcast, half the tile's columns are never written, and the kernel is one block timed once, so the figure it prints is launch overhead. Check that the word indices are distinct before you call an access conflicted, and give the kernel enough work that the clock sees the reads.

Nsight Compute counts the serialization with l1tex__data_bank_conflicts_pipe_lsu_mem_shared_op_ld.sum, and it needs root on a stock driver, where RmProfilingAdminOnly is 1 and a plain run returns ERR_NVGPUCTRPERM. That counter is missing below because day 15 was not run under a profiler, not because the node lacks one: day 11 reports Nsight Compute sector counts from this same card. What follows is clock time against a degree predicted before the run, which is the half a reader on Colab can reproduce anyway.

Measured

Day 15 runs one kernel and passes the stride as a runtime argument, so every row below executes the same instructions in the same order and only the addresses move. Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75.

Word stride Predicted degree Time vs stride 1
1 1 0.286 ms 1.00
2 2 0.563 ms 1.97
4 4 1.117 ms 3.91
32 32 8.886 ms 31.10
33 1 0.286 ms 1.00

gcd(stride, 32) filled the middle column before the run, and every measured ratio lands within 3 percent of it, always a little under. A 31x slowdown out of the fast memory, decided by nothing except which lane read which word.

The same page measures what it is worth inside a kernel anyone would write. A 4096 by 4096 tiled transpose, 128 MiB moved per launch, both variants reading and writing global memory in 128-byte runs:

Kernel Time GB/s Share of the copy ceiling
copy, no transpose 0.567 ms 236.6 100.0%
transpose, tile[32][32] 0.994 ms 135.1 57.1%
transpose, tile[32][33] 0.688 ms 195.2 82.5%

Clearing the conflict buys 45 percent more throughput, 135.1 GB/s to 195.2, and 25 points of the copy ceiling. That is a long way short of 31x, and nothing is inconsistent there: the transpose also waits on global memory, so fixing the shared read moves the total by the share the shared read owned.

Code

From code/day15-bank-conflicts/bank_conflicts.cu. The degree is computed at compile time, so the claim this term rests on cannot rot without turning the build red.

constexpr int conflictDegree(int stride) {
    int a = stride % kBanks;
    int b = kBanks;
    while (a != 0) {
        const int t = b % a;
        b = a;
        a = t;
    }
    return b;
}

static_assert(conflictDegree(kTileDim) == kBanks, "a 32-wide tile column is 32-way");
static_assert(conflictDegree(kTileDim + 1) == 1, "one column of padding clears it");

Diagram

bank-conflicts, preset stride-sweep: 32 lane markers wired to 32 bank slots, with the stride on a control. At stride 1 every wire ends somewhere different; at stride 4 the wires bundle four deep into eight slots; at stride 32 all 32 end in one slot and the request splits into 32 passes.

Alt text: "Thirty-two lanes wired to thirty-two banks at three strides. Stride one fills every bank once, stride four stacks four lanes per bank in eight banks, stride thirty-two stacks all thirty-two in one."

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. This entry stays a draft until a named author and a different named reviewer sign it.