← Glossary
CUDA glossaryPrecision
CC 7.5

What is FP16 in CUDA?

A 16-bit floating-point format with five exponent bits and ten stored mantissa bits, giving compact storage but a maximum finite value of 65,504.

FP16, exposed by CUDA as __half, uses half the storage of FP32. That can halve matrix traffic and shared-memory footprint, but the small exponent field makes range the first trap: a finite FP32 value can become infinity during conversion. Its ten stored mantissa bits give much finer spacing than BF16, so FP16 is often more accurate while values stay in range.

Storage type does not have to be the arithmetic type. A useful FP16 matrix multiply reads half values, converts or multiplies them in hardware, and keeps the running sum in a float accumulator. That is mixed precision: narrow inputs save bytes while FP32 accumulation avoids rounding every partial sum back to ten mantissa bits. Tensor-core APIs such as WMMA make this input/accumulator pairing explicit.

The common mistake is to treat “FP16” as one accuracy number. Range, input rounding and accumulation are separate budgets. Check for non-finite values before judging relative error, and compare against a reference built from the original FP32 inputs, not the already-rounded half copies.

Measured

On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 71 measured __float2half(65520.0f) as infinity and 300000.0f as infinity. In its 2048-square matmul, FP16 storage with FP32 accumulation had maximum absolute error 2.016e-03 and took 23.2897 ms, versus 27.8516 ms for the FP32-storage kernel. The FP16-accumulate variant reached 5.078e-01 error, 252x the FP32-accumulate error. All values come from the 2026-09-02 transcript.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.