← Glossary
CUDA glossaryEcosystem
CC 7.5

What are the free GPU tiers?

Colab, Kaggle and Lightning give a GPU at no cost, with different cards, quotas and traps.

You do not need to own a GPU to write CUDA: 92 of this course's 101 days run free on a Tesla T4. What separates the tiers is not speed. It is what each one refuses to do.

Tier Card Quota The trap
Compiler Explorer Tesla T4, 7.5 20 seconds to compile, 20 seconds to run, no SLA Only the nvcc entries execute, 49 of the 153 CUDA compilers listed. NVRTC and clang entries compile and stop
Colab, free usually a T4, not guaranteed 12 hour sessions, the rest unpublished Profiler counters are restricted
Kaggle, default Tesla P100, 6.0 30 GPU hours a week, 12 hour sessions, 20 GB disk A current toolkit cannot compile for it at all
Kaggle, T4 x2 two T4s the same The only free two-GPU tier, which days 91 and 92 need
Lightning T4, L4, A100 the pricing page says up to 80 free GPU hours Our earlier note recorded 30 credits, roughly 75 T4 hours. The two disagree; check before you plan around it
Modal any, metered $30 of credit a month Billed for load time and a 60 second idle window after your last input, not only for the kernel

The Kaggle row is the one that stops people cold. The free default is a Tesla P100, which is compute capability 6.0, and CUDA 13 removed offline compilation below Turing. So nvcc never gets to your code: nvcc fatal : Unsupported gpu architecture 'sm_60'. Nothing is misconfigured and nothing is missing. Change the accelerator dropdown to T4 x2 and the same file builds. That dropdown is also what gets you two GPUs, which no other free tier offers, so it is the setting to use even on single-GPU days.

The second trap is the profiler, and it is worse because it appears late. Nsight Compute needs GPU performance counters, and without them it prints ERR_NVGPUCTRPERM. On a machine where you have root the fix is small: plain ncu fails, sudo ncu works, and the blocker is the stock driver default RmProfilingAdminOnly: 1 in /proc/driver/nvidia/params. Colab and Kaggle give you neither root nor a way to flip that parameter, which is why every profiling lesson here ships its own .ncu-rep and .nsys-rep so the exercise still works from a report.

Pick by what you are doing. A snippet in a browser goes to Compiler Explorer. A lesson with a harness and a few hundred megabytes goes to Colab or Kaggle. Anything needing two cards goes to Kaggle T4 x2. On a Mac, none of this changes: Metal is a different machine and no emulator is worth the afternoon.

Measured

On a Tesla T4 (driver 595.84, CUDA 12.6, V12.6.85, built with nvcc -O3 -arch=sm_75), day 3 reported 1 CUDA device visible, compute capability 7.5, 14912 MiB of global memory, a toolkit at 12.6 and a driver that accepts up to 13.2. Captured 2026-08-30 on the project's own T4, which is the same card model Colab free and Kaggle T4 x2 hand out, not a capture from either service. Two numbers there are worth carrying: nvidia-smi calls the card 15360 MiB while the runtime reports 14912 MiB, so budget from the smaller one, and one visible device is what you get everywhere except T4 x2. The ERR_NVGPUCTRPERM result above was measured the same way while writing day 11, on a T4 with driver 595.84. Full run behind How to set up CUDA.

Related terms

Compiler Explorer · Metal and the Mac · Nsight Compute · compute capability · NVLink · nvcc

Where you meet this

Day 3, how to set up CUDA, routes you to a tier and hands you the program that prints the card; day 0 is the hardware decision itself. The per-tier detail lives on Colab and Kaggle, and learn CUDA without a GPU is the front door when you have no card at all. The Kaggle failure has its own page: nvcc fatal: Unsupported gpu architecture.

Sources

Byline

Author and reviewer are unassigned. This entry publishes when two different named people have signed it; the written and last-checked dates are set then.