← Glossary
CUDA glossaryPrecision
CC 8.0 for tensor-core mma

What is BF16 in CUDA?

A 16-bit float with FP32's eight exponent bits and seven stored mantissa bits, preserving range while giving up precision near ordinary values.

BF16, or bfloat16, uses the same exponent width as FP32 and truncates the precision budget to fit sixteen bits. That makes overflow much less likely than with FP16, whose five exponent bits stop at 65,504. The trade is coarse spacing: around 1.0, BF16 steps are eight times wider than FP16 steps.

That split is useful for model activations and gradients whose dynamic range matters more than another three mantissa bits. It does not make BF16 automatically safer. Small updates can round away, and a running sum still needs a wider accumulator type. In a mixed-precision pipeline, BF16 inputs commonly accumulate into FP32.

The other trap is confusing a data type with an available fast path. CUDA can convert BF16 values on a Turing card, but BF16 tensor-core mma requires compute capability 8.0 or newer. Always separate “the source compiled” from “this GPU supports the instruction.”

Measured

On a Tesla T4 (compute capability 7.5, driver 580.173.02, CUDA 12.6), day 71 converted 300000.0f to 299008.0 in __nv_bfloat16; FP16 converted the same input to infinity. The precision example went the other way: 1.0007 became 1.0000000000 in BF16 and 1.0009765625 in FP16. The program's capability report printed BF16 tensor-core support as no on this card, with sm_80 as the floor. These are conversion and support results, not a BF16 GEMM timing.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.