← Glossary
CUDA glossaryPrecision
CC 7.5

What is FMA in CUDA?

Fused multiply-add: one instruction that computes a*b+c with a single rounding, which is both faster and more accurate than doing it in two steps.

Done separately, a multiply rounds its product and an add rounds again, two chances to lose bits. The fused instruction keeps the full product internally and rounds once at the end. NVIDIA's floating point whitepaper walks the classic case (https://developer.nvidia.com/sites/default/files/akamai/cuda/files/NVIDIA-CUDA-Floating-Point.pdf , checked 2026-08-30): when a*b and c nearly cancel, the answer lives entirely in the bits the intermediate rounding throws away, so the two-step version returns zero and the fused version returns the exact result. More accurate and faster is the rare free lunch, which is why the compiler contracts a*b + c into FMA by default and why every CUDA core throughput figure assumes it.

The confusion this entry exists to kill: because contraction is controlled by -fmad, which -use_fast_math also sets, people file FMA under "fast math, therefore sloppy". It is the opposite. Turning it off with -fmad=false costs accuracy in cancellation-prone code as well as speed. What FMA does change is bit-for-bit agreement with a CPU that evaluated in two roundings, which matters for reproducibility comparisons, not for closeness to the true answer.

The contraction is a build property, not a callsite one, and day 47 fingerprints it at runtime: a probe computes x*x + y and checks whether the fused and unfused forms return the same bits, so every table in that lesson states which compiler behavior actually produced it.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), captured 2026-09-01. Day 47 computes x*x + y over 1,048,576 elements with y chosen so the exact answer is a tiny power of two and never zero, then counts instructions with sudo ncu:

kernel max rel err FFMA FMUL FADD
residualFused 0.000e+00 1048576 0 0
residualSeparate 1.000e+00 0 1048576 1048576

The error column is the whitepaper's story reproduced on this card: the fused kernel is exact on every element, the separate one returns zero everywhere (relative error 1.000e+00), and the instruction counts prove the only difference is one FFMA against an FMUL-FADD pair per element. The times were 0.0511 against 0.0510 ms, indistinguishable at this size because the kernel is memory-bound; the accuracy difference is total. Rebuilt with -fmad=false, the fused kernel's error becomes 1.000e+00 too, confirming the flag, not the source spelling, decides.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.