SETUP / DAY 3

How to set up CUDA (Linux, Windows, WSL2, Colab, Kaggle)

All setup lessons
Day 3in-technical-review

How to set up CUDA (Linux, Windows, WSL2, Colab, Kaggle)

Kaggle hands you a free GPU. You open a notebook, turn the accelerator on, write hello world, and nvcc says:

nvcc fatal   : Unsupported gpu architecture 'sm_60'

Nothing is broken. Kaggle's free default is a Tesla P100 at compute capability 6.0, and CUDA 13 removed offline compilation below Turing. Pick T4 x2 instead.

Sources: https://www.kaggle.com/docs/efficient-gpu-usage and https://docs.nvidia.com/cuda/archive/13.0.1/cuda-toolkit-release-notes/index.html#deprecated-architectures (both checked 2026-08-29).

Most CUDA setup failures come from version mismatches, not corrupt installs. There are two compatibility checks across three numbers: code cannot require a newer capability than the card provides, and the toolkit or runtime cannot require more than the driver supports.

This page links to the install steps for each machine and gives you a program that prints all three numbers. It also explains the build flag used in the course.

Three numbers and two compatibility gates

What your GPU is. Compute capability is the feature version of the GPU, written 7.5 in the docs and sm_75 on the command line. A Tesla T4 is 7.5, an A100 is 8.0, and an RTX 5090 is 12.0.

The value does not change. Look yours up at https://developer.nvidia.com/cuda-gpus (checked 2026-08-29), or read it from the program below.

What your binary was built for. nvcc puts machine code for named architectures into the executable. If none fits the card, the program can compile, link, and start.

The first kernel launch then fails with no kernel image is available for execution on the device.

What your driver will accept. nvidia-smi prints a CUDA version in its header. That is the highest runtime the installed driver supports, not the toolkit you installed.

CUDA 13.x needs driver 580 or newer, and CUDA 13.3 Update 1 needs 610.43.02 (https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html#cuda-driver , Tables 2 and 3, checked 2026-08-29).

The floor for this course is 7.5. Run nvcc --list-gpu-arch and your toolkit prints exactly which architectures it accepts; on 13.3.0 that list starts at compute_75 and there is nothing below it, because 13.0 dropped offline compilation and library support for Maxwell, Pascal and Volta (https://docs.nvidia.com/cuda/archive/13.0.1/cuda-toolkit-release-notes/index.html#deprecated-architectures , checked 2026-08-29). Which generation your card belongs to decides whether the current toolkit will talk to it at all.

A 7.5 card runs sm_75 code under a CUDA 13-capable driver; compute_80 code fails at launch on that card, while CUDA 13 nvcc rejects sm_60 before a driver is involved.

The one flag, and why it is that one

-arch=sm_75, on every build in this course, even when it looks redundant.

It is redundant in one narrow sense: sm_75 is already nvcc's default (https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html#gpu-architecture-arch , "sm_75 is used as the default value; PTX is generated for compute_75 then assembled and optimized for sm_75", checked 2026-08-29). The default moved once already, from sm_52 in CUDA 11.0, and a build that silently retargets under you is a bad way to find out.

-arch=sm_75 also works on newer cards because it is shorthand. The nvcc manual says --gpu-architecture=sm_86 expands to --gpu-architecture=compute_86 --gpu-code=sm_86,compute_86; the same rule at 75 puts Turing machine code and compute_75 PTX in the binary.

If the machine code does not match the card, the driver compiles the PTX at run time. This is JIT compilation (https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html , sections 4.2.7.2 and 5.7.2.2, checked 2026-08-29).

The binary runs native code on sm_75 and can JIT on newer cards. It will not run on an older card or use newer instructions. Day 69 uses -gencode to include machine code for several architectures.

The intuition to drop: -arch=native is not the safe default. It reads like the right answer. It detects only the GPUs visible on the build machine and generates no PTX at all, so the binary is silently non-portable to every other card, including the one your notebook is given tomorrow.

Where to install what

Find your row. Each one ends at a page that is only about that machine.

Your machine What you install Where
Linux, NVIDIA GPU driver first, then the toolkit https://docs.nvidia.com/cuda/cuda-installation-guide-linux/index.html
Windows, NVIDIA GPU Windows driver, then WSL2, then the toolkit inside WSL Windows and WSL2
macOS, any Mac nothing. There is no CUDA for macOS Learn CUDA without a GPU
AMD, Intel, or integrated graphics nothing the same page
Google Colab nothing, the toolkit is already there Colab
Kaggle nothing, but pick T4 x2, never the default P100 Kaggle
A browser and nothing else nothing Compiler Explorer at https://godbolt.org/

Four more machine problems that would drown this page, each with its own page:

The Mac answer, in one line. CUDA has not run on macOS since CUDA 10.2. The 11.0 release notes state: "CUDA 11.0 does not support macOS for developing and running CUDA applications" (https://docs.nvidia.com/cuda/archive/11.0/cuda-toolkit-release-notes/index.html , Deprecated and Dropped Features, checked 2026-08-29).

Use Metal for local Mac GPU work or a free GPU tier for this course.

The program that reads your card

Full program in code/day03-setup/devicequery.cu. NVIDIA ships a bigger one as deviceQuery in cuda-samples; this is one file so you can paste it into a Colab cell or an editor without cloning anything.

Two rules it follows, and the second one is the interesting half.

Every number comes from the device, not from a table. cudaGetDeviceProperties fills a struct with the card's real values, so the page cannot go stale against your hardware (https://docs.nvidia.com/cuda/cuda-runtime-api/structcudaDeviceProp.html , checked 2026-08-29).

It reports the binary's target next to the card's architecture. These are the two numbers behind the run-time failure. __CUDA_ARCH__ exists only during device compilation, where nvcc assigns it "a three-digit value string xy0 (ending in a literal 0) for each stage 1 nvcc compilation that compiles for compute_xy" (https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html#virtual-architecture-macros , checked 2026-08-29).

A binary built for sm_75 reports 750. The kernel writes that value to device memory, and the host prints it:

__global__ void reportCompiledArch(int* out, size_t n) {
    const size_t i = blockIdx.x * static_cast<size_t>(blockDim.x) + threadIdx.x;
    if (i < n) {
#if defined(__CUDA_ARCH__)
        out[i] = __CUDA_ARCH__;
#else
        out[i] = 0;
#endif
    }
}

That launch is also the test. On a mismatch it is the cudaGetLastError() immediately after the launch that reports error 209, not the cudaDeviceSynchronize() behind it, which is why the program calls both in that order.

Note. One CUDA call in that file is deliberately not wrapped in CUDA_CHECK: the cudaGetDeviceCount at the top. Half the people reading this page do not have a GPU yet, and printing no CUDA-capable device is detected and then quitting is the least useful thing the program could do to them. It names the free tiers instead. Every other call is checked.

Results

Re-verified on a Tesla T4 with driver 580.173.02 and CUDA 13.0 (V13.0.88) on 2026-09-02. The original CUDA 12.6 run under driver 595.84 remains in the page's evidence array; the block below is the CUDA 13.0 transcript.

Your GPU's card
  CUDA devices visible         1
  device 0                     Tesla T4
  compute capability           7.5  (-arch=sm_75)
  this binary was built for    750  (__CUDA_ARCH__)
  runtime version (nvcc)       13.0
  max CUDA the driver takes    13.0
  SMs                          40
  warp size                    32
  max threads per block        1024
  max threads per SM           1024
  global memory                14911 MiB
  shared mem per block         48 KiB
  shared mem per block, opt-in 64 KiB
  L2 cache                     4096 KiB
  32-bit registers per block   65536

Your answers
  BLANK  arch                  work it out from compute capability, written without the dot
  BLANK  warps per SM          work it out from max threads per SM and warp size
  BLANK  max resident threads  work it out from SMs and max threads per SM

Three lines there are worth slowing down on, because they are the version confusion this whole page exists to clear up.

runtime version (nvcc) 13.0 is the toolkit that compiled this binary. max CUDA the driver takes 13.0 is the newest CUDA runtime driver 580.173.02 advertised in this run, and it is what nvidia-smi prints in its CUDA Version: header. Neither is the driver package version.

The older evidence shows the same distinction with toolkit 12.6 and driver 595.84 advertising 13.2. The three numbers have three meanings. The question about the last two has 360,646 views on Stack Overflow (https://stackoverflow.com/questions/53422407, checked 2026-08-30); the NVML mismatch in Pitfalls has even more views.

this binary was built for 750 is __CUDA_ARCH__, the architecture the code was compiled for. It matches the card's 7.5. When those two disagree you get the failure in Pitfalls below.

The three BLANK lines are the exercise. The program will not compute them for you.

The Pitfalls section records the measured no kernel image is available for execution on the device failure. It was reproduced on Compiler Explorer on 2026-08-29 by building with -arch=sm_80 and reading cudaGetLastError() after the launch. The call returned error 209 with that exact string.

The shape to expect:

Line What it is
CUDA devices visible how many GPUs this session can see. Kaggle's T4 x2 is the only free tier where it is more than one
device 0 the marketing name of the card
compute capability x.y, plus the -arch=sm_xy you should pass
this binary was built for xy0, from __CUDA_ARCH__
runtime version / driver version the two numbers that make nvcc and nvidia-smi look like they disagree
SMs, warp size, max threads per SM how many streaming multiprocessors the card has, and how much fits on one
memory, shared memory, L2, registers the rest of the card

Then three graded lines, PASS, FAIL or BLANK, one per answer.

The two lines to read together are compute capability and this binary was built for. On a good build they say the same thing in two notations, 7.5 and 750. Any other arrangement is one of the two failures in the diagram above.

Do not memorise the numbers your card prints. Memorise where each one came from. The point of the exercise is that you can regenerate the card on any machine in half a minute.

Run it yourself

Cheapest first. The program allocates four bytes, launches one kernel and exits, so it sits nowhere near Compiler Explorer's 20 second compile and 20 second run caps. Pin the compiler to a specific nvcc entry rather than trunk, and target sm_75 or lower: the runner is a Tesla T4 (FACT-SHEET.md section 4, checked 2026-08-29).

Colab and Kaggle both work and both give you a shell, which Compiler Explorer does not, so use one of them if you also want to run nvidia-smi and nvcc --version beside the program. On Kaggle, pick T4 x2 (https://www.kaggle.com/docs/notebooks , checked 2026-08-29). On a card you own, the build line is the one at the top of the file:

nvcc -std=c++17 -O3 -arch=sm_75 -o devicequery devicequery.cu

Exercise

Build and run devicequery.cu on whatever GPU you can reach. Read the card it prints, work out the three answers at the top of the file, replace the -1s, and rebuild until it prints PASS three times.

Time: 20 to 30 minutes, most of it getting a GPU. Submit: your edited devicequery.cu and the card it printed.

Check: the program grades itself. A blank answer prints BLANK with the fields to combine, but does not fail the run. A wrong answer prints FAIL, your value, the card's value, and the source fields, then exits non-zero.

Three PASS lines and exit code 0 pass the check.

Hint 1

None of the three answers is a line you can copy off the card. Two of them combine two lines, and one of them is a line with the punctuation taken out.

Hint 2

For the two that combine: an SM holds some number of threads, a warp is some number of threads, and the GPU holds some number of SMs. Which pair gives you warps, and which pair gives you the total?

Solution

There is no separate starter file. devicequery.cu ships with the three constants at -1, so the diff is those three lines.

arch is the compute capability without the dot: 7.5 becomes 75, 8.9 becomes 89, and 12.0 becomes 120. That value follows sm_, so -arch=sm_7.5 is invalid.

warps per SM is max threads per SM divided by warp size. max resident threads is SMs times max threads per SM. The Results section shows the values from the tested card.

The third number is the one worth carrying: how many threads can be resident on the whole GPU at once. Every occupancy argument from day 11 onward is a fraction of it.

Pitfalls

nvcc fatal : Unsupported gpu architecture 'sm_60' Your toolkit no longer supports the target you requested. Switch to T4 x2 on Kaggle, or install CUDA 12.x for a card below 7.5. The three spaces before the colon match nvcc's output.

More on this error.

no kernel image is available for execution on the device You built for a newer architecture than your card supports, such as -arch=sm_80 on a compute capability 7.5 card. The program compiles, links, and starts, but the launch fails.

Compiler Explorer reproduced this on 2026-08-29. After a kernel built with -arch=sm_80 launched, cudaGetLastError() returned error 209: no kernel image is available for execution on the device.

More on this error.

nvcc: command not found, but nvidia-smi works. The two commands come from different packages. nvidia-smi ships with the driver, while nvcc ships with the toolkit.

The toolkit may be missing, or /usr/local/cuda/bin may not be on your PATH. More on this error.

Failed to initialize NVML: Driver/library version mismatch A package upgrade replaced the driver files while the old kernel module stayed loaded. Reboot, or unload and reload the modules.

An apt upgrade can cause this mismatch.

More on this error.

Windows has two common setup failures. nvcc from plain PowerShell gives nvcc fatal : Cannot find compiler 'cl.exe' in PATH, because cl.exe is on PATH only inside a Developer Command Prompt (more on this error).

Under WSL2, do not install a Linux display driver. NVIDIA states: "This is the only driver you need to install. Do not install any Linux display driver in WSL" (https://docs.nvidia.com/cuda/wsl-user-guide/index.html , sections 3.1 and 4, checked 2026-08-29).

The guide also warns against the cuda and cuda-drivers meta-packages; installing them can cause no CUDA-capable device is detected.

Go deeper

Next

You have a GPU you can reach and a card that names it. Day 4 spends those numbers on blocks, threads, and which thread touches element 1337. If you arrived here without a GPU, day 0 has the hardware decision tree.