What is FMA in CUDA?
Fused multiply-add: one instruction that computes a*b+c with a single rounding, which is both faster and more accurate than doing it in two steps.
Done separately, a multiply rounds its product and an add rounds again, two chances to lose bits. The fused instruction keeps the full product internally and rounds once at the end. NVIDIA's floating point whitepaper walks the classic case (https://developer.nvidia.com/sites/default/files/akamai/cuda/files/NVIDIA-CUDA-Floating-Point.pdf , checked 2026-08-30): when a*b and c nearly cancel, the answer lives entirely in the bits the intermediate rounding throws away, so the two-step version returns zero and the fused version returns the exact result. More accurate and faster is the rare free lunch, which is why the compiler contracts a*b + c into FMA by default and why every CUDA core throughput figure assumes it.
The confusion this entry exists to kill: because contraction is controlled by -fmad, which -use_fast_math also sets, people file FMA under "fast math, therefore sloppy". It is the opposite. Turning it off with -fmad=false costs accuracy in cancellation-prone code as well as speed. What FMA does change is bit-for-bit agreement with a CPU that evaluated in two roundings, which matters for reproducibility comparisons, not for closeness to the true answer.
The contraction is a build property, not a callsite one, and day 47 fingerprints it at runtime: a probe computes x*x + y and checks whether the fused and unfused forms return the same bits, so every table in that lesson states which compiler behavior actually produced it.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), captured 2026-09-01. Day 47 computes x*x + y over 1,048,576 elements with y chosen so the exact answer is a tiny power of two and never zero, then counts instructions with sudo ncu:
| kernel | max rel err | FFMA | FMUL | FADD |
|---|---|---|---|---|
| residualFused | 0.000e+00 | 1048576 | 0 | 0 |
| residualSeparate | 1.000e+00 | 0 | 1048576 | 1048576 |
The error column is the whitepaper's story reproduced on this card: the fused kernel is exact on every element, the separate one returns zero everywhere (relative error 1.000e+00), and the instruction counts prove the only difference is one FFMA against an FMUL-FADD pair per element. The times were 0.0511 against 0.0510 ms, indistinguishable at this size because the kernel is memory-bound; the accuracy difference is total. Rebuilt with -fmad=false, the fused kernel's error becomes 1.000e+00 too, confirming the flag, not the source spelling, decides.
Related terms
Where you meet this
- Day 47, fast math and precision, the lesson that owns this term and ran the cancellation test.
- Day 49, the measured roofline, whose peak-FLOP ceiling counts each FMA as two operations.
- Day 68, reproducibility, where matching a CPU bit for bit is the goal and
-fmadis a knob.
Sources
- CUDA Programming Guide, mathematical functions appendix, for FMA accuracy and the intrinsic forms: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/mathematical-functions.html (checked 2026-08-30)
- NVIDIA, "Precision and Performance: Floating Point and IEEE 754 Compliance for NVIDIA GPUs", for the single-rounding definition and the cancellation example: https://developer.nvidia.com/sites/default/files/akamai/cuda/files/NVIDIA-CUDA-Floating-Point.pdf (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.