← Glossary
CUDA glossaryLibraries
CC depends on the selected kernel; measured sm_75 path runs at 7.5

What is CUTLASS?

NVIDIA's header-only CUDA C++ template library for composing architecture-specific matrix multiplication and related kernels from explicit data, tile, pipeline and epilogue choices.

CUTLASS exposes the choices that a higher-level library normally makes behind an API call. A GEMM type fixes the input and accumulator types, matrix layouts, threadblock and warp tiles, instruction shape, pipeline stages and epilogue. The resulting CUDA kernel is specialized at compile time for an architecture tag.

CUTLASS does not inherently require Ampere. The architecture tag and kernel policy determine the minimum compute capability. A Turing build can use tensor core mma instructions without Ampere's asynchronous copy pipeline. CuTe is the layout algebra used by newer CUTLASS APIs.

Measured

On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 84 built CUTLASS v4.7.1 commit cb4247394dd82148787aed73e5dc7cef33cbf862 for sm_75 and sm_80. The sm_75 GEMM passed all 262,144 output comparisons, with its worst error at 0.194 of the tolerance budget.

The SASS held 128 HMMA.1688.F32 instructions for sm_75 and 64 HMMA.16816.F32 instructions for sm_80. The sm_80 binary correctly refused to execute on the CC 7.5 T4, so these measurements establish sm_80 build and disassembly, not sm_80 runtime behavior.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified evidence was captured on 2026-09-02; publication still requires named author and reviewer sign-off.