Why are floating point reductions nondeterministic?
Repeatable floating-point results require a fixed input, reduction shape, addition order and compiled instruction sequence.
Floating-point addition is not associative: rounding after (a + b) + c can differ from a + (b + c). NVIDIA's floating-point guide explains why operation order and fused instructions affect the answer (https://docs.nvidia.com/cuda/floating-point/index.html, checked 2026-09-01). A parallel reduction adds scheduling to that problem. When blocks finish with atomicAdd, their arrival order is not fixed, so the final bits may move even when every input is identical.
A fixed tree removes scheduler choice from the addition order. It does not make the answer mathematically exact, and it does not promise the same bits after changing the launch shape, compiler or GPU. Deterministic and accurate are separate claims.
Measured
On a Tesla T4 with driver 595.84 and CUDA 12.6, day 68 ran each reduction 100 times. The ordinary atomic tail produced 6 distinct patterns, Kahan accumulation followed by the same atomic tail produced 5, and the two-pass fixed tree produced 1. Their maximum relative errors were 2.753e-07, 1.864e-07 and 1.739e-09 respectively. Rebuilding with -fmad=false changed the atomic run hashes but left the fixed tree's bits unchanged for this input.
Related terms
Where you meet this
- Day 68, floating point and reproducibility, the 100-run comparison.
- Day 26, atomics, the earlier arrival-order sweep.
- Day 47, fast math, where compilation changes floating-point instructions.
Sources
- NVIDIA Floating Point and IEEE 754, for order, rounding and FMA behavior: https://docs.nvidia.com/cuda/floating-point/index.html (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came from the verification node on 2026-09-01.