← Glossary
CUDA glossaryEcosystem
CC 7.5

What is Compiler Explorer (godbolt)?

The site that compiles and runs CUDA in a browser on a real GPU, which is how every runnable snippet on this site works.

Most people know it as the place you paste C++ to read the assembly, and stop there. It also executes. Asking its API for the CUDA list returns 153 compilers, 49 of which report supportsExecute: true, and those 49 are exactly the NVCC entries, from 9.1.85 through 13.3.0. The ones that only compile are the NVRTC entries, the clang-based CUDA compilers and the SCALE entries. This project's own research file called the whole thing compile-only; a POST to the executor with executorRequest=true came back with exit code 0 and GPU: Tesla T4 sm_75 on stdout, which settled it.

That card is the constraint everything else follows from. The runner is a Tesla T4 at compute capability 7.5, so a snippet built for a newer architecture compiles and then fails at run time with no kernel image is available for execution on the device. Pass -arch=sm_75 or lower and pin the embed to a specific compiler id such as nvcc133, never to trunk, whose output drifts under you. The sandbox is nsjail on g4dn spot instances, with a 20 second compile cap and a 20 second run cap, and the API documentation offers no SLA and reserves the right to rate limit.

What fits inside 20 seconds is one file, one launch, no data files. What does not fit is a benchmark. Day 11's harness moves 536,870,912 bytes for each of seven access patterns with warm-ups and repeats, which is nowhere near a 20 second budget, and it needs a profiler besides. For those, use a free tier or your own card: free GPU tiers covers Colab and Kaggle, and learn CUDA without a GPU is the front door if you own no NVIDIA hardware. Compiler Explorer stays the fastest way to answer "does this compile, and what does nvcc make of it", with the PTX and SASS panes open next to the source.

Measured

The project's verification node is a Tesla T4 at compute capability 7.5 (driver 595.84, CUDA 12.6, nvcc -O3 -arch=sm_75), the same card model Compiler Explorer reports for its runner. Day 3 confirms what that means for a binary: built with -arch=sm_75, the kernel reads __CUDA_ARCH__ as 750, and the same card reports 40 SMs, 48 KiB of shared memory per block and a 4096 KiB L2. A snippet that runs here runs there.

Two numbers say what to expect from an embed. Day 1 ran the same hello world 20 times on each of two T4s: block 0 printed first 16 times on one and 18 on the other, so a reader who reruns an embed can legitimately see the device lines in a different order than the page shows. And profiling is not part of the deal: on day 11 plain ncu on this node fails with ERR_NVGPUCTRPERM and needs root, because the stock driver ships RmProfilingAdminOnly: 1. A browser sandbox gives you no root. Captured 2026-08-30, except day 11 on 2026-08-29.

Related terms

Where you meet this

Sources

Byline

Written by: unassigned. Reviewed by: unassigned. This entry is a draft and cannot publish until both are named people, and two different ones. Written on: not set. Last checked: not set. Numbers captured 2026-08-30 on the project's verification node, except day 11 on 2026-08-29.