← Glossary
CUDA glossaryPrecision
CC 8.0 for 2:4 sparse tensor-core acceleration

2:4 structured sparsity

2:4 structured sparsity keeps exactly two values in each consecutive group of four along a matrix's supported dimension.

The fixed pattern gives hardware enough structure to compress the matrix and skip known zero positions without storing an arbitrary sparse index for every value. On supported NVIDIA tensor cores, libraries such as cuSPARSELt can use the pattern for sparse matrix multiplication. Pruning a dense matrix into the pattern changes its values, so performance and numerical quality must be evaluated separately.

Day 98 applies magnitude pruning: within every four-value group it keeps the two entries with the largest absolute values. The Tesla T4 has compute capability 7.5 and does not provide the Ampere-generation 2:4 sparse tensor-core path, so the lesson records pruning accuracy only. It does not report a cuSPARSELt speedup or imply that zeros alone accelerate the dense CUDA-core kernel.

Measured

At matrix size 2048, the pruned result had maximum absolute error 4.120e+01 and RMS error 2.865. The dense per-column INT8 result on the same workload had maximum error 8.311 and RMS error 0.2468, making the pruned RMS error about 11.6x larger.

No sparse-kernel timing or profiler artifact was captured because the verification GPU lacks the required sparse tensor-core capability. The intended cuSPARSELt speedup comparison therefore remains unmeasured rather than being inferred from the number of zeros.

Diagram: Each consecutive group of four weights retains the two largest magnitudes, while the other two positions become zero before sparse tensor-core multiplication.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.