What is a bank conflict in CUDA shared memory?
Two lanes of a warp hitting different addresses in the same bank, which serializes their accesses.
You write one shared read, and the memory answers it in as many passes as its busiest bank needs. NVIDIA states the rule without hedging: "If multiple addresses of a memory request map to the same memory bank, the accesses are serialized. The hardware splits a memory request that has bank conflicts into as many separate conflict-free requests as necessary, decreasing the effective bandwidth by a factor equal to the number of separate memory requests." So the quantity worth computing is the conflict degree, the largest number of distinct words any single bank is asked for. Degree 1 is one request. Degree 32 is 32 requests for an instruction you issued once.
For the pattern that causes most of them there is a closed form. Lane L reading word L * s lands in bank L * s mod 32; two lanes collide when their index difference times s is a multiple of 32, which happens every 32 / gcd(s, 32) lanes, so each occupied bank collects gcd(s, 32) lanes and their words all differ. The degree is gcd(stride, 32), and the size of the stride has nothing to do with it. That is the opposite of what day 11 taught about global memory, where every step up the stride cost more bandwidth than the last. Shared memory is periodic rather than monotone: strides 31, 33, 63 and 65 are all free, and only the factors a stride shares with 32 are ever charged for.
The usual way to get this wrong is to time a broadcast and report it as a conflict. Lanes reading the same address are "coalesced into a single multicast", not serialized, and the distinction is where a widely copied demo falls over. A published 120-day CUDA challenge indexes its slow kernel with tx * 2 % 32, which puts lanes tx and tx + 16 on one address: the writes race, the reads broadcast, half the tile's columns are never written, and the kernel is one block timed once, so the figure it prints is launch overhead. Check that the word indices are distinct before you call an access conflicted, and give the kernel enough work that the clock sees the reads.
Nsight Compute counts the serialization with l1tex__data_bank_conflicts_pipe_lsu_mem_shared_op_ld.sum, and it needs root on a stock driver, where RmProfilingAdminOnly is 1 and a plain run returns ERR_NVGPUCTRPERM. That counter is missing below because day 15 was not run under a profiler, not because the node lacks one: day 11 reports Nsight Compute sector counts from this same card. What follows is clock time against a degree predicted before the run, which is the half a reader on Colab can reproduce anyway.
Measured
Day 15 runs one kernel and passes the stride as a runtime argument, so every row below executes the same instructions in the same order and only the addresses move. Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75.
| Word stride | Predicted degree | Time | vs stride 1 |
|---|---|---|---|
| 1 | 1 | 0.286 ms | 1.00 |
| 2 | 2 | 0.563 ms | 1.97 |
| 4 | 4 | 1.117 ms | 3.91 |
| 32 | 32 | 8.886 ms | 31.10 |
| 33 | 1 | 0.286 ms | 1.00 |
gcd(stride, 32) filled the middle column before the run, and every measured ratio lands within 3 percent of it, always a little under. A 31x slowdown out of the fast memory, decided by nothing except which lane read which word.
The same page measures what it is worth inside a kernel anyone would write. A 4096 by 4096 tiled transpose, 128 MiB moved per launch, both variants reading and writing global memory in 128-byte runs:
| Kernel | Time | GB/s | Share of the copy ceiling |
|---|---|---|---|
| copy, no transpose | 0.567 ms | 236.6 | 100.0% |
transpose, tile[32][32] |
0.994 ms | 135.1 | 57.1% |
transpose, tile[32][33] |
0.688 ms | 195.2 | 82.5% |
Clearing the conflict buys 45 percent more throughput, 135.1 GB/s to 195.2, and 25 points of the copy ceiling. That is a long way short of 31x, and nothing is inconsistent there: the transpose also waits on global memory, so fixing the shared read moves the total by the share the shared read owned.
Code
From code/day15-bank-conflicts/bank_conflicts.cu. The degree is computed at compile time, so the claim this term rests on cannot rot without turning the build red.
constexpr int conflictDegree(int stride) {
int a = stride % kBanks;
int b = kBanks;
while (a != 0) {
const int t = b % a;
b = a;
a = t;
}
return b;
}
static_assert(conflictDegree(kTileDim) == kBanks, "a 32-wide tile column is 32-way");
static_assert(conflictDegree(kTileDim + 1) == 1, "one column of padding clears it");
Diagram
bank-conflicts, preset stride-sweep: 32 lane markers wired to 32 bank slots, with the stride on a control. At stride 1 every wire ends somewhere different; at stride 4 the wires bundle four deep into eight slots; at stride 32 all 32 end in one slot and the request splits into 32 passes.
Alt text: "Thirty-two lanes wired to thirty-two banks at three strides. Stride one fills every bank once, stride four stacks four lanes per bank in eight banks, stride thirty-two stacks all thirty-two in one."
Related terms
Where you meet this
- Day 13, shared memory, where the tile that conflicts is introduced.
- Day 14,
__syncthreads(), the barrier between the tile store and the tile load. - Day 15, shared memory bank conflicts explained, the lesson that owns this term.
- Day 16, tiled matrix multiply, where the tile shape is chosen with this arithmetic in hand.
- Colab setup, where the timing runs and the profiler counters do not.
Sources
- CUDA C++ Best Practices Guide 10.2.3.1, for the serialization rule and the multicast exception: https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html (checked 2026-08-30)
- CUDA Programming Guide 2.3.4.2.2, "Shared Memory Bank Conflicts": https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/writing-cuda-kernels.html (checked 2026-08-30)
- "What is a bank conflict? (Doing Cuda/OpenCL programming)", 131 votes and 71,084 views: https://stackoverflow.com/questions/3841877/what-is-a-bank-conflict-doing-cuda-opencl-programming (checked 2026-08-29)
- The demo with no conflict in it, quoted above: https://github.com/AdepojuJeremy/CUDA-120-DAYS--CHALLENGE/blob/main/daily-updates/day-12-Bank-Conflicts-in-Shared-Memory.md (checked 2026-08-30)
- NVIDIA on
ERR_NVGPUCTRPERMand performance-counter permissions: https://developer.nvidia.com/nvidia-development-tools-solutions-err_nvgpuctrperm-permission-issue-performance-counters (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. This entry stays a draft until a named author and a different named reviewer sign it.