CUDA, Triton, cuTile, CUTLASS, Mojo, HIP, Metal and WebGPU: what to learn
GPU tools differ in hardware support and in how much of the kernel they generate. Day 84's CUTLASS C++ path supports compute capability 7.5, while its optional Ampere build and day 86's Triton GPU path need 8.0.
Check hardware support first, then decide how much control the tool should take. This lesson compares those properties and applies them to five cases.
What each tool decides for you
The first axis is how much of the machine you describe. CUDA C++ lets you name the block size, the index arithmetic, the guard, the shared-memory layout and the vector width. Day 44 compares four matmul kernels whose only difference is shared loads per FMA, 2.0 then 1.125 then 0.25 then 0.0625, each ratio a decision written into the source.
At the other end, a library call names only the operation. Day 39 found that CUB beat the hand-written reduce by 1.38x and the hand-written scan by 2.33x, and day 40 measured a 5.4x gap on PageRank. The library usually wins.
Kernel-writing experience still helps you inspect its implementation and identify operations that the library does not cover.
Triton and cuTile take a tile shape and mask, then choose the thread schedule. CUTLASS handles instruction selection and pipelining while exposing layout algebra.
HIP keeps a CUDA-like programming model for AMD GPUs. Metal and WebGPU target different machine models.
Diagram: how much of the machine you still describe. One horizontal axis, "you decide" at the left and "the tool decides" at the right, each tool a labelled marker with its required compute capability printed underneath. Band 1, left: CUDA C++, "block size, index, guard, shared layout, vector width". Caption "day 44: shared loads per FMA 2.0 to 0.0625, every ratio hand-written". Band 2, middle: Triton and cuTile, "you pick the tile, the compiler picks the threads". Caption "compute capability 8.0 floor: a T4 at 7.5 runs neither". Band 3, right: cuBLAS, cuDNN and CUB, "you pick the operation". Caption "day 39: CUB 1.38x on reduce, 2.33x on scan". Alt text: "One axis from writing every hardware decision yourself to naming only the operation. CUDA C++ at the left, Triton and cuTile in the middle behind a compute capability 8.0 floor, CUB at the right where it beat hand-written scan by 2.33 times."
The map, and where each row comes from
| Tool | Runs on | You write | It decides | It cannot |
|---|---|---|---|---|
| CUDA C++ | NVIDIA, sm_75 and up on 13.x | C++ with <<< >>> |
register allocation only | leave NVIDIA without a port |
| Triton | NVIDIA CC 8.0+, AMD ROCm 6.2+, Linux only | Python, tile at a time | threads, vectorisation, scheduling | run on a T4, or on a Mac wheel |
| cuTile | NVIDIA CC 8.x through 12.x, driver r580+ | Python or C++ tiles | the same, inside the toolkit | take a tile dimension that is not a compile-time power of two |
| CUTLASS | NVIDIA, Turing through Blackwell on the 2.x C++ API; newer paths vary | C++ templates, or Python CuTe DSL | instruction selection, pipelining | make one API path cover every architecture |
| HIP | AMD, via ROCm | C++ that looks like CUDA | nothing new | be a drop-in replacement for CUDA |
| Mojo | NVIDIA, AMD and Apple silicon, per a vendor matrix | Mojo | a lot, and it is young | show you a third-party maturity record yet |
| Metal | Apple silicon, A14 and M1 and up | Metal shading language | its own memory model | run CUDA, on any Mac, at all |
| WebGPU | browsers, and natively through wgpu | WGSL | portability over hardware access | reach a tensor core |
Each row is tied to a source checked while this page was written.
- Triton: "Supported Platforms: Linux", "Supported Hardware: NVIDIA GPUs (Compute Capability 8.0+)", "AMD GPUs (ROCm 6.2+)" (https://github.com/triton-lang/triton/blob/main/README.md#compatibility , checked 2026-09-01). Release 3.8.0 publishes manylinux wheels and nothing else (https://pypi.org/pypi/triton/json , same date).
- cuTile: "A GPU with compute capability 8.x, 9.x, 10.x, 11.x or 12.x" and "NVIDIA Driver r580 or later" (https://docs.nvidia.com/cuda/cutile-python/quickstart.html); "Tile dimensions must be compile-time constants that are powers of two" (https://docs.nvidia.com/cuda/cutile-python/index.html). Both checked 2026-08-29.
- CUTLASS: the 2.x C++ API used on day 84 supports the T4's Sm75 target. Current CUTLASS also provides "CUDA C++ template abstractions and Python domain-specific languages (DSLs)" for newer architectures, with Hopper and Ampere Python support marked experimental (https://docs.nvidia.com/cutlass/latest/index.html , checked 2026-09-01).
- HIP: "HIP is not intended to be a drop-in replacement for NVIDIA CUDA, and developers should expect to do some manual coding and performance tuning work to port existing projects to AMD GPUs" (https://rocm.docs.amd.com/projects/HIP/en/latest/what_is_hip.html , checked 2026-09-01).
- Mojo: the compatibility page grades hardware "Continuously tested" or "Known compatible", puts one card, the B200, in the first tier and the T4 and Apple M1 through M5 in the second (https://mojolang.org/docs/requirements/ , checked 2026-09-01). Both tiers are Modular's account of Modular's CI.
- Metal: "a modern, tightly integrated graphics and compute API coupled with a powerful shading language designed so you can take full advantage of Apple silicon", on A14 and M1 and later (https://developer.apple.com/metal/ , checked 2026-09-01).
- WebGPU: a W3C Candidate Recommendation Draft dated 20 August 2026 (https://www.w3.org/TR/webgpu/), WGSL's dated 31 August 2026 (https://www.w3.org/TR/WGSL/). MDN grades the API "Limited availability", "not Baseline", secure context only (https://developer.mozilla.org/en-US/docs/Web/API/WebGPU_API). All checked 2026-09-01.
Why the table omits speed
Two commonly repeated numbers fail the source check: a cuTile figure of
1007 TFLOP/s on a B200 with a 2.5x win over
FlashAttention-2, and a Triton claim of 62 to 101 percent of cuBLAS
across three cards. Both trace to one source that predates CUDA 13.3's
Hopper support, and this project's verification pass refuted both
(FACT-SHEET.md section 8). No independent post-13.3 comparison of cuTile
against Triton was found, so this page does not rank their speed.
The table separates control from hardware support. More convenience gives the compiler more decisions, while the hardware requirement remains fixed for a given release.
One kernel, two owners
The repo has the same fused axpy-and-clamp twice, in
code/day89-tool-map/fused_axpy.cu
and
code/day89-tool-map/triton_axpy.py.
Both use 1,000,003 elements, a tile of 1024, the same alpha, and the same
tolerance rule. Neither file measures time.
Their source shows which indexing choices each tool exposes.
Both files check elementwise agreement with a host reference at 4 float epsilons plus an absolute floor. The floor handles exact zeros from the clamp, where relative tolerance does not apply. They also require at least one clamped value, so non-negative test input cannot skip the branch unnoticed.
The CUDA half:
// You own the decomposition here: which thread reads which element, the
// guard that keeps the tail in bounds, and the launch shape that puts
// blockIdx where you want it. The arithmetic is two lines in the middle.
__global__ void fusedAxpyClamp(const float* __restrict__ x,
float* __restrict__ y, float a, size_t n) {
const size_t i = static_cast<size_t>(blockIdx.x) * blockDim.x + threadIdx.x;
if (i < n) {
const float v = a * x[i] + y[i];
y[i] = v > 0.0f ? v : 0.0f;
}
}
The Triton version applies the same arithmetic to the same elements:
@triton.jit
def fused_axpy_clamp(x_ptr, y_ptr, alpha, n_elements,
BLOCK_SIZE: tl.constexpr):
pid = tl.program_id(axis=0)
offsets = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
mask = offsets < n_elements
x = tl.load(x_ptr + offsets, mask=mask)
y = tl.load(y_ptr + offsets, mask=mask)
v = alpha * x + y
tl.store(y_ptr + offsets, tl.maximum(v, 0.0), mask=mask)
offsets is a tile of indices, while i names one element. The Triton file
therefore never names a lane; the README's
grep for threadIdx, blockIdx, blockDim, __syncthreads and
warpSize is the check.
Both end in the same place on NVIDIA hardware.
nvcc emits PTX here, and Triton writes its own PTX
beside its other IRs when TRITON_KERNEL_DUMP is set, with
TRITON_DUMP_DIR naming "the directory to save the dumped IR and
ptx/amdgcn" (same README, checked 2026-09-01).
The pair does not compare speed because neither file records time. The Triton GPU path also requires compute capability 8.0 or newer.
Results
CUDA and Triton interpreter paths verified; Triton GPU remains blocked.
The corrected CUDA program passed on the Tesla T4 on 2026-09-02 at all
1,000,003 elements, with
zero mismatches and 162,896 clamped outputs. Its grid was 977 with 445 tail
lanes, it reported that compute capability 7.5 misses the tile-DSL floor, its
PTX head named .target sm_75, and the explicit-index grep counts were CUDA 2
and Triton 0.
The retained failed attempt counted 153,846 device clamps against 158,371 in
the host reference. One repeated input pair rounded to zero under separate
host operations but to 2^-27 under device FMA; its 4,525 occurrences exactly
explain the difference.
Moving the input off that boundary produced the final
passing CUDA result. The passing and failed transcripts are separate files in
code/day89-tool-map/evidence/. The dated standalone PTX artifact is also
retained there.
A separate Triton 3.7.1 and torch 2.13.0+cu126 environment capture accompanies an interpreter run that passed both gates at all 1,000,003 elements and exited 0.
| Step | Needs a GPU | Expected |
|---|---|---|
nvcc -arch=compute_75 -ptx |
no | passed; .target sm_75 in captured head |
./fused_axpy |
yes, compute capability 7.5 | passed both gates |
grep -c over both files |
no | CUDA 2, Triton 0 |
TRITON_INTERPRET=1 python3 triton_axpy.py |
no | passed both gates, exit 0 |
python3 triton_axpy.py on the T4 |
yes, compute capability 8.0 | blocked by T4 capability 7.5 |
The access snapshot in FACT-SHEET.md section 4 recorded Compiler Explorer,
Colab, and Kaggle as hosted NVIDIA options. It also recorded an isolated
five-second Modal run near $0.011 on a T4 and $0.07 on an H100, against a $30
monthly credit. These prices are retained as dated evidence, not run guidance;
see FACT-SHEET.md before using them.
Four things the run can prove wrong:
- Verified for the emitted PTX. The dump succeeded and its
.targetline names sm_75. The standalone dated artifact isevidence/fused_axpy-2026-09-02.ptx. Anything else would mean I had misread how nvcc spells a virtual architecture in the file it emits, which is the basis for calling PTX the common denominator. - Verified.
./fused_axpypassed both gates, printed a grid of 977 with 445 tail lanes, and reported that this device misses the tile-DSL floor. The grid is arithmetic rather than a measurement: 1000003 over a tile of 1024 rounds up to 977, and 977 tiles cover 445 elements that do not exist. - Verified. The two grep counts are 2 and 0. Two lines of the CUDA file name a thread-level identifier and no line of the Triton file does. A nonzero count on the right kills the claim about who names lanes.
- Partly verified. The interpreter passed on the CPU with zero mismatches and 162896 clamped outputs. The Triton GPU path is still blocked by the T4's compute capability 7.5, so no GPU failure wording or output is claimed.
The evidence directory contains no timing because this comparison does not measure performance.
Run it yourself
The map and PTX build need no GPU. nvcc -arch=compute_75 -ptx writes PTX
without an NVIDIA device.
Running fused_axpy needs compute capability 7.5
or newer, while the Triton GPU path needs 8.0 or newer. Without a supported
GPU, run the Triton interpreter and use the setup options at
/setup/learn-cuda-without-a-gpu.
Exercise
Pick a tool for each of the five situations below and write one sentence defending each pick, naming the axis that decided it. The quiz on this page holds the same five, one question each.
- A fused bias-and-activation epilogue for a training loop that already runs in PyTorch on an A100, wanted this week.
- A GEMM with an unusual epilogue on an H100, expected to land close to cuBLAS.
- A kernel that has to ship to customers running AMD MI300 as well as NVIDIA H100.
- A real-time image filter inside an app that runs on an iPhone, a Mac and in a browser.
- A batched matrix multiply that is 8 percent of your runtime, in a program where cuBLAS already covers the shape.
Time: 25 to 40 minutes. Submit: five tool names and five sentences, each naming one column of the table above.
Check: answer the five quiz questions on the page. Each option explains the relevant axis. No harness grades this exercise.
Hint 1
For four cases, hardware support or an existing library decides the answer before performance. Check those constraints before comparing tools.
Hint 2
Two of the situations name hardware that rules out most of the table. Look at the "Runs on" column and cross out rows before comparing what is left.
Solution
- Triton. The A100 is at compute capability 8.0, so the floor is cleared, the operation is elementwise and suited to fusion, and the kernel lives next to the Python that calls it. Day 86 writes one.
- CUTLASS, on its C++ path. An unusual epilogue on Hopper is what the template library exists for, and the Python DSL is marked experimental below Blackwell.
- Triton again, or two source trees. Its README is the only row in the table that covers both vendors from one source, at ROCm 6.2 and above. HIP is the other answer and the heavier one, because HIPIFY translates part of the way and AMD says plainly that HIP is not a drop-in replacement.
- Metal on Apple, WebGPU in the browser, or WebGPU alone if one code path is worth more to you than the last few percent. CUDA runs on none of those three machines.
- None of them. Call cuBLAS. At 8 percent of runtime a kernel twice as fast buys 4 percent, and day 39 measured the library winning even where the hand-written version was the exercise.
Check hardware first, then library coverage, then the amount of machine detail you need to control. Choose the tool after those constraints.
Pitfalls
You pick Triton for a GPU below compute capability 8.0. Triton's README
documents compute capability 8.0 and above. The interpreter,
TRITON_INTERPRET=1, runs correctness checks on the CPU but provides no GPU
timings.
You run hipify over a kernel full of intrinsics and expect a build. HIPIFY's own documentation says "HIP is not a complete replacement for CUDA, and HIPIFY cannot automatically translate all code", that libraries with no HIP equivalent cannot be translated, and that NVIDIA-tuned code "might require additional rework to optimize performance on AMD GPUs" (https://rocm.docs.amd.com/projects/HIPIFY/en/latest/index.html , checked 2026-09-01). A clean translation starts the port; it does not finish it.
You expect Apple silicon to run CUDA through a compatibility layer. ZLUDA and SCALE do not provide that path. Use Metal or WebGPU for local Apple silicon, or run CUDA on NVIDIA hardware.
You use a DSL without checking its generated code. A profiler can expose work the DSL cannot express, and its reports still refer to PTX and SASS. Day 46 explains how to inspect them.
You compare languages and omit CCCL. The fastest correct answer on days 39 and 40 was a library call, not a language change, and tiling a kernel by hand in newer syntax does not beat an algorithm someone already tuned.
OpenCL, SYCL and OpenACC are also options. OpenCL 3.1 is the portable floor with the weakest tensor-core story (https://www.khronos.org/opencl/); SYCL 2020 is the pick when the requirement is one modern ISO C++ source across three vendors (https://www.khronos.org/sycl/); OpenACC is "a collection of compiler directives" for offloading loops in existing C, C++ and Fortran (https://www.openacc.org/specification). All three checked 2026-09-01.
Lifetime Stack Overflow tag counts, which measure attention rather than merit: cuda 14,762, opencl 5,777, openacc 408, sycl 178 (https://api.stackexchange.com/2.3/tags/opencl%3Bcuda%3Bsycl%3Bopenacc/info?site=stackoverflow , checked 2026-09-01). Mojo is younger than all of them and every claim about it here is Modular's own.
Go deeper
- Triton README, "Compatibility" and the environment variables list: https://github.com/triton-lang/triton/blob/main/README.md#compatibility (checked 2026-09-01)
- CUTLASS documentation, the C++ and Python quick starts: https://docs.nvidia.com/cutlass/latest/index.html (checked 2026-09-01)
- HIPIFY documentation, on what it does and does not translate: https://rocm.docs.amd.com/projects/HIPIFY/en/latest/index.html (checked 2026-09-01)
- MDN, WebGPU API, for the browser support table and WGSL: https://developer.mozilla.org/en-US/docs/Web/API/WebGPU_API (checked 2026-09-01)
Next
Day 90 applies this map to the four capstones and checks which ones should have used a library call.