What is INT8 quantization in CUDA?
Representing real values as signed eight-bit integers plus scale and optional zero point, then accumulating integer products and rescaling the output.
Symmetric INT8 quantization chooses a scale such as max_abs / 127, rounds each real value divided by that scale, and stores the result from -127 to 127. A matrix multiply accumulates integer products, usually in INT32, then multiplies by the A and B scales. An asymmetric activation adds a zero point and subtracts its contribution in the epilogue.
Scale granularity decides which values share a range budget. One scale for a weight matrix lets outlier columns set the step for every other column. Per-column scales preserve smaller channels at the price of one scale load per output column. MX block scaling applies the same idea to shorter blocks and floating formats.
INT8 storage is not proof that tensor cores ran. Day 98's T4 supports INT8 tensor cores, but its teaching kernel performs integer MADs on CUDA cores. The transcript names that mechanism so its modest speedup is not attributed to hardware it never invoked.
Measured
On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 98 measured size 2048. FP32 took 32.2266 ms. One-scale INT8 took 23.6386 ms, with maximum error 7.333 and RMS error 0.4899. Per-column INT8 took 23.7686 ms, maximum error 8.311 and RMS error 0.2468, a 1.36x speedup over FP32. Its worst statistical column fraction improved 11.57x over one scale. The deterministic K-scaled correctness gate passed for every per-column output.
Related terms
Where you meet this
- Day 98, quantization and sparsity, which produced the measurements above.
- Day 66, testing CUDA kernels, where the error gate is derived.
- Colab setup, whose T4 supports INT8 tensor cores even though this kernel does not use them.
Sources
- CUDA compute-capability appendix, tensor-core input support: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/compute-capabilities.html (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.