What are FP32 and FP64 in CUDA?
The 32-bit single-precision and 64-bit double-precision IEEE formats, used respectively for ordinary GPU arithmetic and higher-accuracy references or scientific work.
FP32 is CUDA C++'s float: eight exponent bits and 23 stored mantissa bits, with four bytes per value. FP64 is double: eleven exponent bits, 52 stored mantissa bits and eight bytes per value. FP64 offers far finer resolution and wider range, but doubles storage and traffic. Its hardware throughput also varies sharply by GPU class, so precision choice and performance choice cannot be separated from the device.
For kernel testing, FP64 is often most valuable on the CPU side. A double-precision reference accumulates more accurately than the FP32 kernel it judges and can expose rounding error without sharing the kernel's summation behavior. That is different from demanding an FP64 GPU kernel. Day 71 uses original FP32 inputs, a host double reference, and an FP32 CUDA baseline before comparing FP16 storage and alternate accumulator types.
Do not assume “double is twice as accurate” or infer an FP64 throughput ratio from an FP32 timing. Accuracy depends on the operation count and conditioning; throughput needs its own measured kernel. The verified owner run below measures FP32 against an FP64 reference, not FP64 GPU speed, so this page does not invent the inventory's still-missing direct ratio.
Measured
On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 71 ran the 2048-square FP32 tiled matmul in 27.8516 ms. Against the host FP64 reference, maximum absolute error was 7.610e-05 and the K-scaled gate fraction was 0.1952. At 256 and 1024, errors were 4.582e-06 and 1.740e-05. The transcript contains no FP64 GPU throughput measurement, so no FP64:FP32 speed ratio is claimed here.
Related terms
Where you meet this
- Day 71, mixed precision, which produced the FP32-versus-FP64-reference results above.
- Day 49, roofline, where arithmetic throughput is measured explicitly.
- Learn CUDA without a GPU, where reference reasoning remains possible without device timing.
Sources
- CUDA C++ Programming Guide, floating-point support and arithmetic behavior: https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.