How to set up CUDA (Linux, Windows, WSL2, Colab, Kaggle)
Kaggle hands you a free GPU. You open a notebook, turn the accelerator on, write hello world, and nvcc says:
nvcc fatal : Unsupported gpu architecture 'sm_60'
Nothing is broken. Kaggle's free default is a Tesla P100 at compute capability 6.0, and CUDA 13 removed offline compilation below Turing. Pick T4 x2 instead.
Sources: https://www.kaggle.com/docs/efficient-gpu-usage and https://docs.nvidia.com/cuda/archive/13.0.1/cuda-toolkit-release-notes/index.html#deprecated-architectures (both checked 2026-08-29).
Most CUDA setup failures come from version mismatches, not corrupt installs. There are two compatibility checks across three numbers: code cannot require a newer capability than the card provides, and the toolkit or runtime cannot require more than the driver supports.
This page links to the install steps for each machine and gives you a program that prints all three numbers. It also explains the build flag used in the course.
Three numbers and two compatibility gates
What your GPU is. Compute capability is the feature version of the GPU, written 7.5 in the docs and sm_75 on the command line. A Tesla T4 is 7.5, an A100 is 8.0, and an RTX 5090 is 12.0.
The value does not change. Look yours up at https://developer.nvidia.com/cuda-gpus (checked 2026-08-29), or read it from the program below.
What your binary was built for. nvcc puts machine code for named architectures into the executable. If none fits the card, the program can compile, link, and start.
The first kernel launch then fails with no kernel image is available for execution on the device.
What your driver will accept. nvidia-smi prints a CUDA version in its header. That is the highest runtime the installed driver supports, not the toolkit you installed.
CUDA 13.x needs driver 580 or newer, and CUDA 13.3 Update 1 needs 610.43.02 (https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html#cuda-driver , Tables 2 and 3, checked 2026-08-29).
The floor for this course is 7.5. Run nvcc --list-gpu-arch and your toolkit prints exactly which architectures it accepts; on 13.3.0 that list starts at compute_75 and there is nothing below it, because 13.0 dropped offline compilation and library support for Maxwell, Pascal and Volta (https://docs.nvidia.com/cuda/archive/13.0.1/cuda-toolkit-release-notes/index.html#deprecated-architectures , checked 2026-08-29). Which generation your card belongs to decides whether the current toolkit will talk to it at all.
The one flag, and why it is that one
-arch=sm_75, on every build in this course, even when it looks redundant.
It is redundant in one narrow sense: sm_75 is already nvcc's default (https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html#gpu-architecture-arch , "sm_75 is used as the default value; PTX is generated for compute_75 then assembled and optimized for sm_75", checked 2026-08-29). The default moved once already, from sm_52 in CUDA 11.0, and a build that silently retargets under you is a bad way to find out.
-arch=sm_75 also works on newer cards because it is shorthand. The nvcc manual says --gpu-architecture=sm_86 expands to --gpu-architecture=compute_86 --gpu-code=sm_86,compute_86; the same rule at 75 puts Turing machine code and compute_75 PTX in the binary.
If the machine code does not match the card, the driver compiles the PTX at run time. This is JIT compilation (https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html , sections 4.2.7.2 and 5.7.2.2, checked 2026-08-29).
The binary runs native code on sm_75 and can JIT on newer cards. It will not run on an older card or use newer instructions. Day 69 uses -gencode to include machine code for several architectures.
The intuition to drop: -arch=native is not the safe default. It reads like the right answer. It detects only the GPUs visible on the build machine and generates no PTX at all, so the binary is silently non-portable to every other card, including the one your notebook is given tomorrow.
Where to install what
Find your row. Each one ends at a page that is only about that machine.
| Your machine | What you install | Where |
|---|---|---|
| Linux, NVIDIA GPU | driver first, then the toolkit | https://docs.nvidia.com/cuda/cuda-installation-guide-linux/index.html |
| Windows, NVIDIA GPU | Windows driver, then WSL2, then the toolkit inside WSL | Windows and WSL2 |
| macOS, any Mac | nothing. There is no CUDA for macOS | Learn CUDA without a GPU |
| AMD, Intel, or integrated graphics | nothing | the same page |
| Google Colab | nothing, the toolkit is already there | Colab |
| Kaggle | nothing, but pick T4 x2, never the default P100 | Kaggle |
| A browser and nothing else | nothing | Compiler Explorer at https://godbolt.org/ |
Four more machine problems that would drown this page, each with its own page:
- Check your CUDA version, the most-viewed CUDA question on Stack Overflow.
- nvcc and nvidia-smi disagree, which is the second.
- Which CUDA version works with your driver.
- Picking which GPU your job runs on, and using a GPU from Docker.
The Mac answer, in one line. CUDA has not run on macOS since CUDA 10.2. The 11.0 release notes state: "CUDA 11.0 does not support macOS for developing and running CUDA applications" (https://docs.nvidia.com/cuda/archive/11.0/cuda-toolkit-release-notes/index.html , Deprecated and Dropped Features, checked 2026-08-29).
Use Metal for local Mac GPU work or a free GPU tier for this course.
The program that reads your card
Full program in code/day03-setup/devicequery.cu. NVIDIA ships a bigger one as deviceQuery in cuda-samples; this is one file so you can paste it into a Colab cell or an editor without cloning anything.
Two rules it follows, and the second one is the interesting half.
Every number comes from the device, not from a table. cudaGetDeviceProperties fills a struct with the card's real values, so the page cannot go stale against your hardware (https://docs.nvidia.com/cuda/cuda-runtime-api/structcudaDeviceProp.html , checked 2026-08-29).
It reports the binary's target next to the card's architecture. These are the two numbers behind the run-time failure. __CUDA_ARCH__ exists only during device compilation, where nvcc assigns it "a three-digit value string xy0 (ending in a literal 0) for each stage 1 nvcc compilation that compiles for compute_xy" (https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html#virtual-architecture-macros , checked 2026-08-29).
A binary built for sm_75 reports 750. The kernel writes that value to device memory, and the host prints it:
__global__ void reportCompiledArch(int* out, size_t n) {
const size_t i = blockIdx.x * static_cast<size_t>(blockDim.x) + threadIdx.x;
if (i < n) {
#if defined(__CUDA_ARCH__)
out[i] = __CUDA_ARCH__;
#else
out[i] = 0;
#endif
}
}
That launch is also the test. On a mismatch it is the cudaGetLastError() immediately after the launch that reports error 209, not the cudaDeviceSynchronize() behind it, which is why the program calls both in that order.
Note. One CUDA call in that file is deliberately not wrapped in
CUDA_CHECK: thecudaGetDeviceCountat the top. Half the people reading this page do not have a GPU yet, and printingno CUDA-capable device is detectedand then quitting is the least useful thing the program could do to them. It names the free tiers instead. Every other call is checked.
Results
Re-verified on a Tesla T4 with driver 580.173.02 and CUDA 13.0 (V13.0.88)
on 2026-09-02. The original CUDA 12.6 run under driver 595.84 remains in
the page's evidence array; the block below is the CUDA 13.0 transcript.
Your GPU's card
CUDA devices visible 1
device 0 Tesla T4
compute capability 7.5 (-arch=sm_75)
this binary was built for 750 (__CUDA_ARCH__)
runtime version (nvcc) 13.0
max CUDA the driver takes 13.0
SMs 40
warp size 32
max threads per block 1024
max threads per SM 1024
global memory 14911 MiB
shared mem per block 48 KiB
shared mem per block, opt-in 64 KiB
L2 cache 4096 KiB
32-bit registers per block 65536
Your answers
BLANK arch work it out from compute capability, written without the dot
BLANK warps per SM work it out from max threads per SM and warp size
BLANK max resident threads work it out from SMs and max threads per SM
Three lines there are worth slowing down on, because they are the version confusion this whole page exists to clear up.
runtime version (nvcc) 13.0 is the toolkit that compiled this binary. max CUDA the driver takes 13.0 is the newest CUDA runtime driver 580.173.02
advertised in this run, and it is what nvidia-smi prints in its CUDA Version: header. Neither is the driver package version.
The older evidence shows the same distinction with toolkit 12.6 and driver 595.84 advertising 13.2. The three numbers have three meanings. The question about the last two has 360,646 views on Stack Overflow (https://stackoverflow.com/questions/53422407, checked 2026-08-30); the NVML mismatch in Pitfalls has even more views.
this binary was built for 750 is __CUDA_ARCH__, the architecture the code was compiled for. It matches the card's 7.5. When those two disagree you get the failure in Pitfalls below.
The three BLANK lines are the exercise. The program will not compute them for you.
The Pitfalls section records the measured no kernel image is available for execution on the device failure. It was reproduced on Compiler Explorer on 2026-08-29 by building with -arch=sm_80 and reading cudaGetLastError() after the launch. The call returned error 209 with that exact string.
The shape to expect:
| Line | What it is |
|---|---|
CUDA devices visible |
how many GPUs this session can see. Kaggle's T4 x2 is the only free tier where it is more than one |
device 0 |
the marketing name of the card |
compute capability |
x.y, plus the -arch=sm_xy you should pass |
this binary was built for |
xy0, from __CUDA_ARCH__ |
runtime version / driver version |
the two numbers that make nvcc and nvidia-smi look like they disagree |
SMs, warp size, max threads per SM |
how many streaming multiprocessors the card has, and how much fits on one |
| memory, shared memory, L2, registers | the rest of the card |
Then three graded lines, PASS, FAIL or BLANK, one per answer.
The two lines to read together are compute capability and this binary was built for. On a good build they say the same thing in two notations, 7.5 and 750. Any other arrangement is one of the two failures in the diagram above.
Do not memorise the numbers your card prints. Memorise where each one came from. The point of the exercise is that you can regenerate the card on any machine in half a minute.
Run it yourself
Cheapest first. The program allocates four bytes, launches one kernel and exits, so it sits nowhere near Compiler Explorer's 20 second compile and 20 second run caps. Pin the compiler to a specific nvcc entry rather than trunk, and target sm_75 or lower: the runner is a Tesla T4 (FACT-SHEET.md section 4, checked 2026-08-29).
Colab and Kaggle both work and both give you a shell, which Compiler Explorer does not, so use one of them if you also want to run nvidia-smi and nvcc --version beside the program. On Kaggle, pick T4 x2 (https://www.kaggle.com/docs/notebooks , checked 2026-08-29). On a card you own, the build line is the one at the top of the file:
nvcc -std=c++17 -O3 -arch=sm_75 -o devicequery devicequery.cu
Exercise
Build and run devicequery.cu on whatever GPU you can reach. Read the card it prints, work out the three answers at the top of the file, replace the -1s, and rebuild until it prints PASS three times.
Time: 20 to 30 minutes, most of it getting a GPU. Submit: your edited devicequery.cu and the card it printed.
Check: the program grades itself. A blank answer prints BLANK with the fields to combine, but does not fail the run. A wrong answer prints FAIL, your value, the card's value, and the source fields, then exits non-zero.
Three PASS lines and exit code 0 pass the check.
Hint 1
None of the three answers is a line you can copy off the card. Two of them combine two lines, and one of them is a line with the punctuation taken out.
Hint 2
For the two that combine: an SM holds some number of threads, a warp is some number of threads, and the GPU holds some number of SMs. Which pair gives you warps, and which pair gives you the total?
Solution
There is no separate starter file. devicequery.cu ships with the three constants at -1, so the diff is those three lines.
arch is the compute capability without the dot: 7.5 becomes 75, 8.9 becomes 89, and 12.0 becomes 120. That value follows sm_, so -arch=sm_7.5 is invalid.
warps per SM is max threads per SM divided by warp size. max resident threads is SMs times max threads per SM. The Results section shows the values from the tested card.
The third number is the one worth carrying: how many threads can be resident on the whole GPU at once. Every occupancy argument from day 11 onward is a fraction of it.
Pitfalls
nvcc fatal : Unsupported gpu architecture 'sm_60' Your toolkit no longer supports the target you requested. Switch to T4 x2 on Kaggle, or install CUDA 12.x for a card below 7.5. The three spaces before the colon match nvcc's output.
no kernel image is available for execution on the device You built for a newer architecture than your card supports, such as -arch=sm_80 on a compute capability 7.5 card. The program compiles, links, and starts, but the launch fails.
Compiler Explorer reproduced this on 2026-08-29. After a kernel built with -arch=sm_80 launched, cudaGetLastError() returned error 209: no kernel image is available for execution on the device.
nvcc: command not found, but nvidia-smi works. The two commands come from different packages. nvidia-smi ships with the driver, while nvcc ships with the toolkit.
The toolkit may be missing, or /usr/local/cuda/bin may not be on your PATH. More on this error.
Failed to initialize NVML: Driver/library version mismatch A package upgrade replaced the driver files while the old kernel module stayed loaded. Reboot, or unload and reload the modules.
An apt upgrade can cause this mismatch.
Windows has two common setup failures. nvcc from plain PowerShell gives nvcc fatal : Cannot find compiler 'cl.exe' in PATH, because cl.exe is on PATH only inside a Developer Command Prompt (more on this error).
Under WSL2, do not install a Linux display driver. NVIDIA states: "This is the only driver you need to install. Do not install any Linux display driver in WSL" (https://docs.nvidia.com/cuda/wsl-user-guide/index.html , sections 3.1 and 4, checked 2026-08-29).
The guide also warns against the cuda and cuda-drivers meta-packages; installing them can cause no CUDA-capable device is detected.
Go deeper
- nvcc manual, section 5.7.2 "Shorthand" and section 5.7.4 "Virtual Architecture Macros": https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html (checked 2026-08-29)
- CUDA Installation Guide for Linux: https://docs.nvidia.com/cuda/cuda-installation-guide-linux/index.html , and for Windows: https://docs.nvidia.com/cuda/cuda-installation-guide-microsoft-windows/index.html (both checked 2026-08-29)
cuda-samples,cpp/1_Utilities/deviceQuery: https://github.com/NVIDIA/cuda-samples/tree/master/cpp/1_Utilities/deviceQuery (checked 2026-08-29)- Programming Massively Parallel Processors, 4th edition, chapter 2, on the compilation model: https://shop.elsevier.com/books/programming-massively-parallel-processors/hwu/978-0-323-91231-0
Next
You have a GPU you can reach and a card that names it. Day 4 spends those numbers on blocks, threads, and which thread touches element 1337. If you arrived here without a GPU, day 0 has the hardware decision tree.