What is a tensor core in CUDA?
A specialized arithmetic unit that performs a small matrix multiply-accumulate cooperatively for a warp, at throughput ordinary scalar CUDA-core instructions cannot match.
A CUDA core executes scalar arithmetic for lanes of a warp. A tensor core consumes matrix fragments distributed across a warp and updates an accumulator tile in one instruction. You reach it through an API such as WMMA, explicit PTX such as mma.sync, or a library that selects the instruction for you.
The hardware does not make data movement free. Tiles still have to arrive from global memory and often pass through shared memory; fragment layouts and synchronization still have to be correct. A simple WMMA kernel can therefore use tensor-core arithmetic and remain much slower than cuBLAS. The specialized unit only accelerates the multiply-accumulate portion.
Input support also depends on compute capability. The T4 path accepts FP16 inputs and FP32 accumulation. BF16 and TF32 begin at sm_80, and FP8 mma begins at sm_89. “Has tensor cores” is incomplete unless the format and instruction shape are named.
Measured
On a Tesla T4 (40 SMs, driver 580.173.02, CUDA 12.6), day 72 measured 2048-square FP16-input, FP32-accumulate GEMMs after three warm-ups. The staged WMMA kernel took 7.908 ms and reached 2,172.6 GFLOP/s. The cuBLAS FP16 path took 0.544 ms and reached 31,581.9 GFLOP/s. Both used the same numerical input/accumulator pairing and matched the CPU reference, so the 14.5x throughput gap prices staging and kernel engineering, not a different precision contract.
Diagram: one warp feeds A and B fragments into a tensor-core matrix multiply, then receives an accumulator fragment. The takeaway is that the warp owns the operation collectively; no lane owns a complete tile.
Related terms
Where you meet this
- Day 72, WMMA, which produced the throughput comparison above.
- Day 73, mma.sync, where the instruction and lane ownership become explicit.
- Select which GPU to use, because format support changes with compute capability.
Sources
- CUDA C++ Programming Guide, warp matrix functions: https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html (checked 2026-09-01)
- CUDA compute-capability appendix, tensor-core input support: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/compute-capabilities.html (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.