← Glossary
CUDA glossaryPrecision
CC 7.5

What is mixed precision in CUDA?

Computing with narrow input or storage formats while retaining a wider accumulator, so bandwidth and arithmetic throughput improve without paying narrow-format error at every sum.

Mixed precision makes two decisions separately: how values are stored and multiplied, and how their running sum is represented. FP16 inputs halve bytes relative to FP32 and can feed tensor cores. Keeping the accumulator in FP32 prevents every partial sum from being rounded back to FP16.

This matters most in reductions such as matrix multiplication. Input conversion introduces one rounding error per stored value. Those errors often partly cancel across K products. A narrow accumulator adds another rounding at the scale of the growing partial sum on every iteration, so its error can grow much faster. Wider accumulation cannot recover bits already lost from the input, but it removes that second source.

The term does not mean “turn on half precision everywhere.” BF16, TF32 and FP8 make different range and precision trades and have different hardware floors. A correct experiment names input type, accumulator type, output type and compute capability. It also checks non-finite values before applying a tolerance.

Measured

On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 71 ran identical tiled matmuls with FP16 storage and either FP32 or FP16 accumulation. At K=256, maximum errors were 2.595e-03 and 2.950e-02, a ratio of 11.4. At K=1024 the ratio reached 51.3. At K=2048, FP32 accumulation produced 2.016e-03 maximum error while FP16 accumulation produced 5.078e-01, a 252x ratio. Both paths passed their stated gates; the ratio reports the accuracy price of the accumulator rather than declaring the lower-precision path incorrect.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.