← Glossary
CUDA glossaryPrecision
CC 8.9 for tensor-core mma

What is FP8 in CUDA?

A family of eight-bit floating-point formats that trade range against precision, commonly using E4M3 for activations and E5M2 where wider range matters.

FP8 is not one encoding. E4M3 spends four bits on exponent and three on mantissa; E5M2 shifts one bit from precision to range. E4M3 reaches a largest finite value of 448 and has a gap of 0.125 at 1.0. E5M2 reaches 57,344 with a 0.25 gap. Both are far coarser than FP16, so scaling and a wider accumulator type are essential parts of the computation.

The formats are useful because tensor cores can turn narrow inputs into high matrix throughput. They do not imply that any CUDA GPU can run the instruction. E4M3 and E5M2 mma require compute capability 8.9 or newer; a T4 at 7.5 and an A100 at 8.0 are both below that floor. Mixed precision always includes this hardware question as well as the numerical one.

Another common mistake is quoting a maximum value or bit table as a measured model result. Day 71 only verifies the support gate on its T4. It does not execute an FP8 GEMM, so this entry makes no speed or accuracy claim from that card.

Measured

On a Tesla T4 (compute capability 7.5, driver 580.173.02, CUDA 12.6), day 71 printed “FP8 E4M3/E5M2 (mma, sm_89) no on this card.” That named-device result is the gate: the course's free-tier floor cannot produce an FP8 tensor-core timing. The format table's 448, 57,344, 0.125 and 0.25 values are derived from the documented bit layouts, not represented as benchmark output.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified support evidence was captured on 2026-09-02; publication still requires named author and reviewer sign-off.