← Glossary
CUDA glossaryPrecision
CC 8.0

What is TF32 in CUDA?

A tensor-core input format with FP32's range and at least ten bits of precision, used by Ampere-or-newer matrix instructions rather than ordinary CUDA-core arithmetic.

TF32 exists to put FP32-range data through tensor cores at reduced input precision. It is not a C++ storage type you use for arrays, and it is not the arithmetic ordinary CUDA cores perform for float. The PTX ISA calls its internal layout implementation-defined and promises FP32 range with reduced precision of at least ten bits.

Libraries can select TF32 for a matrix operation whose inputs and outputs are FP32. That is why an “FP32” library benchmark can have different error and speed from a hand-written FP32 kernel. The compute type or math mode is part of the experiment. FP32 accumulation can still be used, but it cannot restore input bits discarded before the multiply; accumulator type and input type are separate controls.

TF32 requires compute capability 8.0 or newer. A T4 can compile host code that discusses the mode, but its sm_75 tensor cores cannot execute TF32 mma. A support line is therefore more honest than a zero or borrowed benchmark.

Measured

On a Tesla T4 (compute capability 7.5, driver 580.173.02, CUDA 12.6), day 71 printed “TF32 (mma, sm_80) no on this card.” The same report printed FP16 support as yes, proving the check distinguished formats rather than merely rejecting tensor cores. No TF32 timing or error was produced because this hardware cannot execute the path.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified support evidence was captured on 2026-09-02; publication still requires named author and reviewer sign-off.