What does -use_fast_math do in CUDA?
The -use_fast_math flag, which swaps precise math functions for faster approximations and turns denormal support off.
It is not one thing. The nvcc manual says the flag implies --ftz=true --prec-div=false --prec-sqrt=false --fmad=true, and on top of those four it substitutes intrinsics for a list of math calls, __expf for expf, __fdividef for division, and so on (https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html , checked 2026-08-30). Each piece has its own cost, documented per function in the math appendix's error tables (https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/mathematical-functions.html , checked 2026-08-30), so the honest question is never "is fast math ok" but which of the five pieces you can afford. Note the bundle includes -fmad=true, which is already the default and improves accuracy; FMA is the piece that does not belong in the flag's reputation.
Two pieces deserve their own paragraph. --ftz=true flushes subnormal results to zero, which is invisible until your data has a tail near the bottom of the float range and then silently zeroes it. And the substitutions are a build property, not a callsite choice: the flag rewrites your expf calls behind your back, so a benchmark of "the same code, two builds" needs a runtime fingerprint to know which functions actually ran, which is exactly what day 47's probe kernel does.
The measured surprise is how little the whole bundle bought on a memory-bound kernel: on day 47's softmax the fast build moved the time by a few percent while roughly doubling the error, because __expf only pays when the kernel is compute-bound and math throughput is the wall. Buy the pieces one at a time, and only after the profiler says the math pipeline is the limit.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), four builds of one source file, captured 2026-09-01. Day 47's softmax (2048 rows of 1024), max relative error against a double reference, mean of 10 timed runs:
| build | exp used | max rel err | time (ms) |
|---|---|---|---|
| default | expf | 7.406e-07 | 0.1073 |
| default | __expf (explicit) | 1.494e-06 | 0.1110 |
| -use_fast_math | all variants | 1.512e-06 | 0.1167 |
Under -use_fast_math all four source spellings collapse to the same 1.512e-06 error, proof the flag rewrote the precise calls. The --ftz=true half is priced by a census of 1024 exponentials ramping into the subnormal range: the default build returns 17 normal, 832 subnormal and 175 zero results with smallest nonzero 1.401298e-45; the ftz build returns 0 subnormals and 1007 zeros with smallest nonzero 1.195105e-38. The 832 values in between did not get less precise; they became zero.
Related terms
Where you meet this
- Day 47, fast math and precision, the lesson that owns this term and priced each piece.
- Day 68, reproducibility, where the substitutions are one more reason two builds disagree.
- Day 97, online softmax, the kernel family where the exp path's speed actually matters.
Sources
- NVCC compiler driver documentation, the
--use_fast_mathsection, for the four implied flags and the substitution list: https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html (checked 2026-08-30) - CUDA Programming Guide, mathematical functions appendix, for the per-function error bounds of the intrinsics: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/mathematical-functions.html (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.