← Glossary
CUDA glossaryPrecision
CC 7.0; lesson path uses sm_75

What is WMMA in CUDA?

CUDA's warp-level matrix API, where 32 threads cooperatively load opaque fragments, execute a matrix multiply-accumulate, and store the resulting tile.

WMMA, the Warp Matrix Multiply-Accumulate API in mma.h, gives a CUDA C++ route to tensor cores. A warp declares matrix A, matrix B and accumulator fragments, fills or loads them, calls mma_sync, then stores the result. No single lane owns the matrix; all 32 lanes participate in every operation.

Fragments are intentionally opaque. You may index the fragment's local storage for elementwise operations that apply uniformly, but the mapping from matrix coordinates to lanes is unspecified and can change with architecture. If an epilogue needs coordinates or a custom layout, mma.sync exposes the lower-level contract. For portable tile arithmetic, WMMA is the safer interface.

Using the API is not enough to approach library throughput. Global loads, shared-memory staging, layout, synchronization and overlap decide whether the tensor core is fed. Day 72's two kernels differ only in whether tiles are staged, yet both remain far behind cuBLAS at large size.

Measured

On a Tesla T4 (40 SMs, driver 580.173.02, CUDA 12.6), day 72 measured the FP16-input, FP32-accumulate staged WMMA kernel at size 2048: 7.908 ms and 2,172.6 GFLOP/s. The direct-global WMMA version took 11.320 ms and reached 1,517.6 GFLOP/s. cuBLAS with the same input/accumulator contract reached 31,581.9 GFLOP/s. The companion layout program printed a complete 16x16 output after touching fragment elements through the API; it demonstrates the opaque storage order without treating it as a portable lane map.

Diagram: a 16x16x16 WMMA operation split across 32 lanes, with fragment boxes labelled opaque. The takeaway is that only the matrix layout is portable, not which lane holds an element.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.