← Glossary
CUDA glossaryPrecision
CC 7.5 for m16n8k8 FP16

What is mma.sync in CUDA PTX?

A PTX warp-level matrix instruction whose shape, layouts and data types determine exactly which registers each lane supplies and receives.

mma.sync is the explicit instruction beneath higher-level tensor-core APIs. Its qualifiers name the tile shape, A and B layouts, destination type, input types and accumulator type. For m16n8k8.row.col.f32.f16.f16.f32, one warp updates a 16x8 FP32 tile from FP16 A and B fragments with K=8.

Unlike WMMA, the PTX contract publishes which fragment values belong in each lane's registers. That enables custom swizzles and coordinate-aware epilogues, but it also makes layout mistakes your responsibility. ldmatrix is an alias on this page because its purpose is to move shared-memory tiles directly into those register layouts.

Instruction availability is shape-specific. FP16 m16n8k8 and ldmatrix require sm_75; m16n8k16 requires sm_80. A T4 can therefore run the useful Turing path many tutorials skip, but source containing an unguarded newer shape fails during assembly even if the host would never launch it.

Measured

On a Tesla T4 (compute capability 7.5, driver 580.173.02, CUDA 12.6), day 73 decoded the A and B registers and verified all 192 fragment elements against the PTX ownership tables. Its m16n8k8 tile and full n=2048 matmul matched double references exactly. The matmul took 2.472 ms and reached 6,951.1 GFLOP/s with FP16 inputs and FP32 accumulation. The m16n8k16 probe printed a supported skip, naming sm_80 as its minimum.

Diagram: lanes 0 through 31 mapped to the A and B coordinates for m16n8k8. The takeaway is that fragment ownership is a published register contract, not a matrix distributed one row per thread.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.