← Glossary
CUDA glossaryPrecision
CC 7.5

What is an accumulator type in CUDA?

The data type used for a running sum or matrix output fragment, chosen independently from input storage to control rounding, range and hardware behavior.

A dot product reads many input pairs but produces one running sum. The accumulator type is the format that sum occupies after each multiply-add. It may be wider than the inputs: FP16 matrices commonly accumulate into FP32. This pairing is the numerical core of mixed precision.

Wider accumulation reduces repeated rounding and raises the partial sum's overflow ceiling. It cannot restore precision discarded when inputs became FP16, but it prevents that storage error from being compounded at every addition. APIs expose the choice differently. WMMA makes accumulator fragment type a template parameter; mma.sync spells input, output and accumulator types in the instruction; cuBLAS uses a compute type.

The performance effect is not a correctness rule. A narrow accumulator can be faster, slower or tied depending on shape and implementation, while still passing a suitable tolerance. Comparisons must keep input, output, geometry and reference fixed and report the accumulator as part of the path name.

Measured

On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 72 built the same WMMA source twice by changing using AccT = float to using AccT = __half. At size 2048, the staged FP32-accumulator kernel took 7.908 ms at 2,172.6 GFLOP/s; the FP16-accumulator exercise took 7.168 ms at 2,396.7 GFLOP/s. Register use changed from 54 to 50 per thread. Both runs reported all four paths matching the CPU reference for the deliberately exact test inputs, so these timings do not prove equal accuracy on arbitrary data.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.