← Glossary
CUDA glossaryMemory
CC 7.5

What are padding and swizzling in CUDA shared memory?

Two fixes for bank conflicts: add a column so rows land on different banks, or XOR the index so the mapping rotates.

Both attack one number, gcd(rowLength, 32), and they differ in what they spend to move it. Padding lengthens the row. Declare tile[32][33] where you had tile[32][32] and row r now begins at word 33r, so a walk down a column steps 33 words at a time, 33 shares no factor with 32, and the 32 lanes land in 32 different banks. Nothing has been aligned and no gap has been left for the hardware to skip. The row length stopped being a multiple of the bank count, and that is the whole mechanism. NVIDIA's transpose post is where most people first copy the + 1, usually without that sentence attached.

The extra column is not free, and it bills in three places. A 32 by 33 float tile is 4224 bytes against 4096, so the T4's 64 KiB of shared memory per SM holds 15 of them where it held 16, which is also the card's ceiling on resident blocks. 33 is not a power of two, so the compiler can no longer fold the row multiply into a shift and every shared access pays an integer add. And on a kernel already close to the per-block limit, one column can cost a resident block and more than the conflict did. Day 15's kernel escapes all of it: at 256 threads per block, the 1024 resident threads per SM cap it at 4 blocks long before shared memory has a say. Work out which limit binds on your kernel before assuming the same.

Swizzling leaves the row 32 wide and permutes inside it, storing element (r, c) at column c ^ r. A row read still covers 32 distinct columns, so it stays conflict free; a column read now takes column c ^ r at row r, which is a different bank on every row. CUTLASS generalises this to a bit-field operation, XOR-ing one field of the offset into another, and describes the result in its own header as AA = ZZ xor YY. The swizzle costs no memory and one XOR of address arithmetic, and it leaves the row a power of two wide so the index stays a shift. Both of those matter to the tensor-core paths from day 73 on, where the tiles are large enough that an extra column per row is memory the block does not have. Day 73 is where this course measures a swizzle. There is no number for one on this page, and the section below says so rather than borrowing the padded kernel's.

Measured

Two kernels one character apart, on a Tesla T4 with driver 595.84 and CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75. Day 15 transposed a 4096 by 4096 float matrix with each of them and timed a straight copy over the same buffers first, as the ceiling. All three moved 128 MiB and all three read and write global memory in 128-byte runs.

Kernel Time GB/s Share of the copy
copy, no transpose 0.567 ms 236.6 100.0%
transpose, tile[32][32] 0.994 ms 135.1 57.1%
transpose, tile[32][33] 0.688 ms 195.2 82.5%

128 bytes per block bought 60 GB/s, and it did so while carrying the handicaps in the paragraph above, which is the direction that cannot flatter the result. On the isolated read probe from the same run the identical change is worth far more: stride 32 costs 8.886 ms and stride 33 costs 0.286, a factor of 31. A transpose keeps a fraction of that because it also waits on global memory. Take the 31x as what the fix is worth to the instruction and the 0.69x as what it was worth to the kernel.

The swizzle was not measured here. Its cost argument is arithmetic, not a benchmark, and the entry will carry a number when day 73 produces one.

Diagram

bank-conflicts, preset pad-33: the same 32 by 32 tile shaded by bank, drawn three times. Unpadded, one column is a single colour. Padded to 33, each row shifts one bank right and the column becomes a diagonal through all 32 colours. Swizzled at c ^ r, the row stays 32 wide and the colours scramble inside it.

Alt text: "One shared tile, three layouts. Unpadded, a column read hits one bank thirty-two times. Padded to thirty-three words, the column steps one bank per row. Swizzled by XOR, the row length is unchanged and the column still covers thirty-two banks."

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. This entry stays a draft until a named author and a different named reviewer sign it.