MX block scaling
MX block scaling assigns one scale to each small block of low-precision values instead of using one scale for an entire tensor.
The shared scale maps each block's largest magnitude into the representable range of a compact element format such as FP8. Smaller blocks adapt more closely to local value ranges, so outliers in one part of a tensor do not reduce precision everywhere else. The scale metadata costs storage and must be applied when values are consumed.
Day 98 models the central tradeoff with groups of 32 values. It is a storage-accuracy census performed on the host, not a benchmark of native MX instructions. The Tesla T4 used for verification predates hardware MX block-scaling support, so the result establishes the quantization effect without claiming an accelerated MX kernel.
The project's checked hardware map places FP4/NVFP4 support on B200 at compute capability 10.0 and exposes it through mma.sync on RTX 50 at compute capability 12.0. That map describes available arithmetic formats; it does not turn the host-side day 98 census into a native MX benchmark.
Measured
For the day 98 input, FP8 E4M3 with one tensor-wide scale had maximum normalized storage error 3.571e-02 and RMS error 1.316e-03. Changing to one scale per 32 values reduced those figures to 7.303e-03 and 8.061e-04. The maximum error fell by 4.89x.
The same census reported per-column INT8 at 3.936e-03 maximum and 9.731e-04 RMS normalized storage error. These are representation errors before matrix multiplication; they are not GEMM output errors or throughput measurements.
Diagram: Two rows of 32 FP8 values share either one global scale or two block-local scales, keeping the second row's small values distinguishable.
Related terms
Where you meet this
- Day 98, quantization and sparsity, which compares tensor-wide, per-column and block-local scaling while separating storage error from GEMM output error.
- Day 71, mixed precision, which establishes the FP8 range that scaling must use.
Sources
- NVIDIA PTX ISA, alternate floating-point data formats: https://docs.nvidia.com/cuda/parallel-thread-execution/#alternate-floating-point-data-formats (checked 2026-09-01)
- CUDA compute-capability appendix, tensor-core input support: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/compute-capabilities.html (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.