Day 0Module 0
in-technical-review

What you need to know before learning CUDA

A learner opened an issue against a published CUDA course on their first day, before writing a single kernel:

test.cu(22): error: a value of type "void *" cannot be assigned to an entity of type "int *"
      ptr = malloc(sizeof(int));
          ^

That report is at https://github.com/Infatoshi/cuda-course/issues/5 (checked 2026-08-29). The line is correct C. It appears in the course's own pointer example.

It fails anyway, and the page does not explain why. You may think your install is broken.

It is not your install. It is that .cu files are compiled as C++, and C++ does not convert void * on its own. This page is the eight things that error is an example of, plus the hardware question you have to answer before day 1: what can you actually run, and what does it cost if the answer is nothing.

Why nvcc rejects code that C accepts

nvcc is not a C compiler with GPU extensions bolted on. It splits your file in two, hands the CPU half to your system compiler and compiles the GPU half itself, and both halves go through a C++ front end.

NVIDIA says so twice on the same page: "The GPU code is implemented as a collection of functions in a language that is essentially C++", and "The default C++ dialect depends on the host compiler. nvcc matches the default C++ dialect that the host compiler uses" (https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html , checked 2026-08-29).

Half of the surprises below follow from that sentence. The rest come from CUDA's two memory spaces. The CPU and its RAM are the host; the GPU and its RAM are the device.

A pointer from malloc and a pointer from cudaMalloc have the same type, the same size and the same printed form, and dereferencing the wrong one on the wrong side can crash with no useful message.

Day 5 uses both kinds of pointer.

You do not need C++ for any of this. You need C, plus the handful of places where nvcc's C++ front end will not accept what your C book taught you. That is a much smaller task than learning C++ first.

A kernel is a function you mark __global__ and launch across thousands of threads; it is not an object, it does not throw, and it has no standard library.

You do not have to own an NVIDIA GPU

Hardware stops many people before day 1, often because they think any fast GPU can run CUDA. A high school student wrote: "I have zero experience with coding in general, not even with languages like Python or C++" (https://www.reddit.com/r/CUDA/comments/1leil5s/ , checked 2026-08-29). The card he had just bought was an RX 9070 XT, which cannot run CUDA.

Neither can a Mac: there is no NVIDIA driver for Apple silicon and there will not be one. Metal covers the Mac path.

Start by asking the machine rather than guessing:

nvidia-smi --query-gpu=name,compute_cap --format=csv

If the command is missing or reports no device, you are on branch A. If your driver rejects compute_cap, drop it and look the card up at https://developer.nvidia.com/cuda-gpus (checked 2026-08-29). The --query-gpu option is documented at https://docs.nvidia.com/deploy/nvidia-smi/index.html (checked 2026-08-29).

Without a usable CUDA 13 GPU, 93 days still run free, 1 is compile-only, and 7 need borrowed or metered hardware; Hopper covers all 101.

Eight of the 101 days need more than compute capability 7.5: days 58, 74, 75, 76, 77, 78, 79 and 86. Six need no GPU at all. The other 87 run on a Tesla T4, which is what all three free tiers hand you.

What you have Runs where you are Read instead of run Closing the gap
Nothing NVIDIA: Mac, AMD, Intel, integrated 93 days, split across Compiler Explorer, Colab and Kaggle's T4 x2 day 78: Compiler Explorer compiles an sm_90a target and the exercise is an annotation 7 days: 58, 74, 75, 76, 77, 79, 86
NVIDIA below CC 7.5 0 days. CUDA 13 removed offline compilation below Turing same as the row above same as the row above, or install CUDA 12.x and accept that the pages will not match your compiler
CC 7.5 exactly: T4, RTX 20 93 days across the local card and Kaggle's free two-GPU tier for 91 and 92 day 78 7 days, same list
CC 8.0 to 8.9: RTX 30, RTX 40, L4, A100 97 days across the local card and Kaggle for 91 and 92 day 78 3 days: 58, 76, 77
CC 9.0 Hopper: H100, H200 all 101; 91 and 92 still want a second card nothing nothing
CC 10.0 or 12.0 Blackwell: B200, RTX 50 100 days; 91 and 92 still want a second card day 78's wgmma, which exists on no Blackwell nothing

The three 9.0 days are each designed to finish inside 15 minutes, and Modal lists H100 SXM5 at $0.001097 per second, which is $3.95 an hour (https://modal.com/pricing , checked 2026-08-29). Three quarter-hour runs are $2.96 of GPU time, but Modal also bills load time and a default 60 second idle keep-alive after your last input (https://modal.com/docs/guide/cold-start , checked 2026-08-29), so budget under $5.

The four days that need 8.0 want an A100 or an L4, and this course carries no rate for either card. Read it off https://modal.com/pricing yourself. An estimate here would be a made-up number in the one place a reader acts on one.

Hardware. Kaggle's default free GPU is a P100 at compute capability 6.0, and CUDA 13 removed offline compilation below sm_75, so the default fails before it starts (https://www.kaggle.com/docs/efficient-gpu-usage and https://docs.nvidia.com/cuda/archive/13.0.1/cuda-toolkit-release-notes/index.html#deprecated-architectures , both checked 2026-08-29). Always pick T4 x2. It is also the only free two-GPU tier, which is what days 91 and 92 need.

If your answer was branch A or branch B, go to learn CUDA without a GPU after this page. If it was C or D, how to install CUDA is the router.

Eight things to check before day 1

The two programs are in code/day00-cuda-prerequisites/, and they are the only code in this course that needs no GPU and no toolkit. Two rules they follow, because the point of a primer is to be checkable:

Every claim on this page is a check in the program, not a sentence. prereq.c prints ok or FAIL and a value for each of ten facts, and returns a non-zero exit code if any of them is wrong on your platform. A primer you cannot run is a primer you skim.

The file that is supposed to fail is a separate file. void_star.c has to compile as C and fail as C++, which is the opposite of what prereq.c needs, so the two cannot share a translation unit.

1. & and *. These two symbols work on both sides of an assignment, and everything else on this page is a variation on them.

    // 1. & takes an address. * reads or writes the thing at that address.
    int x = 41;
    int* p = &x;
    *p = 42;
    check("1. *p = 42 makes x", x, 42);
    check("1. *p reads back", *p, 42);

CUDA cannot avoid them: cudaMalloc hands you an address in the GPU's memory, and passing it to a kernel is the only thing you can do with it.

2. malloc and free, and their device twins. malloc gives you bytes and an address, free gives them back, and every malloc gets exactly one free.

cudaMalloc and cudaFree are that pair for device memory.

    // 2. malloc hands back an address and no type. The cast is required in
    // C++, which is what nvcc compiles a .cu file as, so write it everywhere
    // and the habit costs you nothing. cudaMalloc and cudaFree are the same
    // pair with the memory on the other side of the PCIe bus.
    int* heap = (int*)malloc(4 * sizeof(int));
    if (heap == NULL) {
        fprintf(stderr, "malloc failed\n");
        return EXIT_FAILURE;
    }
    for (int k = 0; k < 4; ++k) {
        heap[k] = 10 * (k + 1);
    }
    check("2. sum of the malloc'd block", sumArray(heap, 4), 100);
    free(heap);

3. The cast nvcc wants and C did not. This is the error at the top of the page.

In C a void * converts to any object pointer type on its own; C++ does not allow that conversion.

nvcc compiles C++.

int* ptr;
ptr = malloc(sizeof(int));   // fine in C, an error under nvcc

The fix is (int*)malloc(sizeof(int)). The same rule is why every CUDA program you will read writes cudaMalloc((void**)&d_a, bytes): &d_a has type float**, the parameter has type void**, and neither language converts between those for you. Full page: a value of type "void *" cannot be assigned.

4. Parenthesize every macro parameter. One learner lost five hours to this: "I am debugging this thing for 5 hours and going nuts" (https://www.reddit.com/r/CUDA/comments/1sidzlf/ , checked 2026-08-29).

A macro substitutes text. It does not evaluate arguments.

// INDEX_BAD is deliberately unparenthesized, so item 4 can show what it
// expands to. A macro substitutes text and does not evaluate its arguments,
// so INDEX_BAD(1 + 1, 0, 4) becomes 1 + 1 * 4 + 0, which is 5, and the author
// wanted 8. Bracket every parameter and the whole body, or write a function.
#define INDEX_BAD(row, col, cols) (row * cols + col)
#define INDEX_OK(row, col, cols) ((row) * (cols) + (col))

Better than brackets: write a function. A __device__ __forceinline__ index function costs nothing at run time and cannot do this to you. The same shape of bug returns in a tiled kernel on day 13 and day 16.

5. Arrays decay to pointers. An array passed to a function becomes a pointer to its first element and forgets how long it was (https://en.cppreference.com/w/c/language/array , checked 2026-08-29).

// The parameter is written as an array and is a pointer. sizeof measures the
// pointer, 8 bytes on any 64-bit machine, not the 40 bytes the caller passed.
// This is why every kernel in this course takes a length beside its pointer:
// there is no other way for the callee to know.
static size_t sizeofParameter(int a[]) {
    return sizeof(a);
}

So void addKernel(float* a, size_t n) is not redundant, and neither is any other kernel signature in this course. Your compiler warns on that sizeof under -Wall, and the warning is right.

6. No host standard library inside a kernel. std::vector, std::string, std::map and <iostream> are compiled for the CPU and cannot be called from device code.

__global__ void k() {
    std::vector<int> v;   // error
    v.push_back(1);
}

NVIDIA names the replacement: "it is recommended to avoid calling a function in the Standard C++ headers std:: from device code ... Instead, it is strongly suggested to call the equivalent functionality in the CUDA C++ Standard Library libcu++, in the cuda::std:: namespace" (https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/cpp-language-support.html , checked 2026-08-29).

libcu++ arrives on day 39. For now, do not use std:: in device code.

The same appendix rules out exceptions and RTTI in device code, so no try, no throw, no dynamic_cast. Full page: calling a host function from a device function.

7. static inside device code is one variable for the whole grid. This one is legal, so nothing warns you.

__device__ void f() {
    static int counter;   // one instance, in global memory
    counter++;            // every thread in the grid racing on it
}

NVIDIA's worked example annotates exactly this as "CORRECT, implicit __device__ memory space specifier" (same appendix, section 5.3.10.4.4, checked 2026-08-29). The implicit __device__ is the trap: one instance per device, not one per thread.

Someone had to spell it out to a learner whose first CUDA program lost to his CPU: "You will have many thousands of threads all simultaneously writing to the same variable" (https://www.reddit.com/r/CUDA/comments/fujwsw/ , checked 2026-08-29). Drop the static and use a plain local, which lives in a register per thread.

8. Compile, then link, and device code has an extra rule. Two commands and what each one produces:

nvcc -arch=sm_75 -c main.cu -o main.o     # compile: source to object
nvcc main.o -o app                        # link: objects to executable

-arch matters more than any other flag. sm_75 is nvcc's current default ("sm_75 is used as the default value", same nvcc page, checked 2026-08-29), which is also why this course's floor is 7.5, and passing a value your toolkit dropped gives you nvcc fatal : Unsupported gpu architecture 'sm_60'.

Device code uses whole-program compilation by default: "the device code cannot reference an entity from a separate file". A __device__ function in b.cu is not callable from a.cu until you pass -rdc=true. That is separate compilation, and CUDA with CMake is the day that does it properly.

Results

The original host-C run used gcc -std=c11 -Wall -Wextra on 2026-08-30. On 2026-09-02, the same checks were re-run on a Tesla T4 with driver 580.173.02 and CUDA 13.0 (V13.0.88). The C prerequisite program passed, and both intended C++ failures occurred under g++ and nvcc.

The original CUDA 12.6 record remains in the evidence list. The prerequisite program needs no GPU; you can run it on the machine you are reading this on.

Day 0 prerequisites, host C only.

ok   1. *p = 42 makes x                 42
ok   1. *p reads back                   42
ok   2. sum of the malloc'd block       100
ok   4. INDEX_BAD(1 + 1, 0, 4)          5
ok   4. INDEX_OK(1 + 1, 0, 4)           8
ok   5. sizeof(a) here                  40
ok   5. sizeof(a) inside a function     8
ok   5. *(a + 2)                        2
ok   5. a[2]                            2
ok   5. (a + 2) - a                     2

0 of 10 checks failed.

Every line is a claim the page makes, checked against what your compiler actually does. Item 5 is the one to stare at: sizeof(a) is 40 in the caller and 8 inside the function, on the same array, because the parameter is a pointer. That is the whole reason every kernel in this course takes a length beside its pointer.

The build is not silent, and that is deliberate:

code/day00-cuda-prerequisites/prereq.c: In function 'sizeofParameter':
code/day00-cuda-prerequisites/prereq.c:51:18: warning: 'sizeof' on array function parameter 'a' will return size of 'int *' [-Wsizeof-array-argument]
   51 |     return sizeof(a);
      |                  ^
code/day00-cuda-prerequisites/prereq.c:50:35: note: declared here
   50 | static size_t sizeofParameter(int a[]) {
      |                               ~~~~^~~

Exactly one warning, and it is the compiler telling you item 5 in its own words. Do not fix it. CI asserts that this warning is present and that there is only one of it, so if a future compiler stops emitting it, the check fails and the lesson gets updated rather than quietly going stale.

What to expect, so you can tell a pass from a failure before then:

Program Command Expected shape
prereq.c gcc -std=c11 -Wall -Wextra -o prereq prereq.c compiles with exactly one warning, on the sizeof in item 5
prereq ./prereq ten ok lines, 0 of 10 checks failed., exit code 0
void_star.c as C gcc -std=c11 -o void_star void_star.c compiles clean, prints 42
void_star.c as C++ g++ -x c++ -std=c++17 -c void_star.c fails on the assignment, exit code 1
void_star.c as CUDA C++ nvcc -std=c++17 -arch=sm_75 -c void_star.cu fails on the assignment, exit code 2 in the captured run

The CUDA 13 transcript matches every row: ten host-C checks passed, the C void * program printed 42, g++ rejected the implicit conversion with exit code 1, and nvcc rejected the same conversion with exit code 2.

Two of those ten values are worth reading rather than counting. Item 4 prints 5 for the unbracketed macro where its author wanted 8. Item 5 prints 40 for sizeof in main and 8 for the same array inside a function.

The labels on each line are this page's item numbers, so they read straight across; item 3 has no line because it is the file that does not compile.

The wording of a compiler error is not portable, and the cause is. GCC, Clang and nvcc all reject the void * assignment and all three phrase it differently. Search on what you did, not on the sentence you got.

Run it yourself

There is no GPU on this page, so there is nothing to rent. Any C compiler works: GCC, Clang, Apple's gcc, or MSVC with the flags adjusted. On Windows without a toolchain, the two files paste into Compiler Explorer as plain C and run there for nothing.

Budget 20 minutes for the two programs and 20 for the check below. If you have no NVIDIA GPU, add learn CUDA without a GPU before day 1; it ends by handing you a working notebook, and this page's branch A is the reason you need it.

Exercise

Ten questions on the eight items above. Answer each one before you open its reveal, then read every explanation, right or wrong, because the explanation is where the teaching is.

Time: 20 minutes. Submit: nothing. Count your score and use the scale at the end.

Check: every answer is on this page, and each reveal names the documentation line or the learner report that settles it. No gate, no score bar; a wrong answer costs nothing.

Q1. int x = 41; int* p = &x; *p = 42; What is x now? (a) 41, (b) 42, (c) 42 only if p is not NULL, (d) undefined.

Answer

(b) 42. *p is the object at that address, and that object is x. (a) reads *p as a copy.

(c) is wrong in an interesting way: p came from &x, which is never null.

(d) needs something out of bounds, and nothing is.

Q2. Predict the output.

int a[4] = {10, 20, 30, 40};
printf("%d\n", *(a + 2));
Answer

30. Pointer arithmetic counts elements, not bytes, so a + 2 advances by two ints. *(a + 2) and a[2] are the same expression written two ways.

Q3. These are lines 21 and 22 of a .cu file. What does nvcc do? (a) compiles it, (b) fails, (c) compiles with a warning, (d) fails because malloc is not available in CUDA.

int* ptr;
ptr = malloc(sizeof(int));
Answer

(b) fails, with error: a value of type "void *" cannot be assigned to an entity of type "int *". (a) is what most C programmers pick, because the line is valid C. (c) is wrong: an error, not a warning.

(d) is wrong twice, since host malloc is fine in host code and device code has its own. Reported at https://github.com/Infatoshi/cuda-course/issues/5 , checked 2026-08-29.

Q4. Predict the output.

#define INDEX(row, col, cols) (row * cols + col)
printf("%d\n", INDEX(1 + 1, 0, 4));
Answer

5. The macro substitutes text, so this expands to 1 + 1 * 4 + 0. The author wanted 8.

Bracket every parameter and the whole body, or write a function. The learner who spent five hours on this exact macro: https://www.reddit.com/r/CUDA/comments/1sidzlf/ , checked 2026-08-29.

Q5. On a 64-bit machine, what prints? (a) 40, (b) 8, (c) 10, (d) it does not compile.

void f(int a[]) { printf("%zu\n", sizeof(a)); }
int main(void) { int a[10]; f(a); }
Answer

(b) 8. The parameter is a pointer, and the array decayed and lost its length. (a) is sizeof in main.

(c) confuses bytes with elements. (d) is why this bites people: it compiles, with a warning most people skip.

Q6. __global__ void k() { std::vector<int> v; v.push_back(1); } (a) it compiles and allocates on the device heap, (b) it compiles but is slow, (c) it fails to compile, (d) it works from compute capability 9.0.

Answer

(c) it fails, with error : calling a host function("std::vector<int, std::allocator<int> > ::push_back") from a __device__/__global__ function("k") is not allowed. (d) is the tempting wrong answer: this has nothing to do with the GPU, it is a compile-time rule.

Use cuda::std:: instead. 56 votes on https://stackoverflow.com/questions/10375680/using-stdvector-in-cuda-device-code , checked 2026-08-29.

Q7. A __device__ function contains static int counter; and 1,024 threads each run counter++. How many counter variables exist? (a) 1,024, (b) 32, (c) one, (d) it does not compile.

Answer

(c) one, in global memory, shared by every thread in the grid. (a) is the assumption that produces the bug. (b) is wrong: nothing gives a warp its own copy of a function-scope static.

(d) is the problem, not the answer, because "static local variables are allowed in device functions" (https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/cpp-language-support.html section 5.3.10.4.4, checked 2026-08-29).

Q8. Predict the output.

float* d_a;
cudaMalloc((void**)&d_a, 4 * sizeof(float));
d_a[0] = 1.0f;
printf("done\n");
Answer

Nothing. It crashes on the assignment, before the printf. cudaMalloc returns an address in the GPU's memory and the CPU cannot dereference it; move data with cudaMemcpy (https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__MEMORY.html , checked 2026-08-29).

The one-line version a learner got back: "cudaMalloc allocates memory on the GPU, and your code is accessing it from the CPU" (https://www.reddit.com/r/CUDA/comments/1eov1et/ , checked 2026-08-29).

Q9. Why cudaMalloc((void**)&d_a, bytes) and not d_a = cudaMalloc(bytes)? (a) C has no return values, (b) the return value is already the error code, so the address goes into a variable you pass by address, (c) device pointers are 128 bits, (d) the cast names the element type.

Answer

(b), and the cast is the second half of it: float** does not convert to void** on its own. (a) is wrong: cudaMalloc uses its return value for cudaError_t.

(c) is wrong: device pointers are ordinary pointers. (d) is wrong, cudaMalloc deals in bytes. Signature at https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__MEMORY.html , checked 2026-08-29.

Q10. a.cu calls a __device__ function defined in b.cu. You run nvcc a.cu b.cu -o app. (a) it works, (b) it works but the function is inlined everywhere, (c) it fails, (d) it works only above compute capability 8.0.

Answer

(c) it fails. "CUDA programs are compiled in the whole program compilation mode by default, i.e., the device code cannot reference an entity from a separate file" (https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html , checked 2026-08-29).

Add -rdc=true and let nvcc device-link. (a) is the intuition ordinary C++ gives you. (d) is wrong: this works on every architecture.

Where to start. Fewer than seven right: read the eight items again and take days 1 to 5 slowly. Seven or eight: start at day 1.

Nine or ten: skim this page and start at day 1 anyway, because questions 3, 7 and 10 are CUDA-specific and no amount of C++ prepares you for them.

Pitfalls

Your first .cu file rejects a line straight out of a C book. The cause is that nvcc compiles .cu as C++, and the most common shape is the uncast malloc: error: a value of type "void *" cannot be assigned to an entity of type "int *". Add the cast.

The full error page ranks the other causes.

You reach for std::vector inside a kernel. error : calling a host function("std::vector<int, std::allocator<int> > ::push_back") from a __device__/__global__ function("k") is not allowed.

Pass a raw pointer and a length instead. Day 39 introduces the device-side library that does have containers.

You buy or borrow a GPU that cannot compile CUDA 13. nvcc fatal : Unsupported gpu architecture 'sm_60', with three spaces before the colon, because that is what nvcc prints and that is what you would paste into a search box.

CUDA 13 removed offline compilation below Turing, so a P100, a V100 or a GTX 1080 is a dead end for this course. Full page: nvcc fatal: Unsupported gpu architecture. How to install CUDA covers the runtime twin of it, no kernel image is available for execution on the device.

You dereference a device pointer on the host. No compiler error, no CUDA error code, just a segmentation fault on a line that looks like ordinary array access. If a crash makes no sense, check which side of the bus the pointer came from first.

You decide to learn C++ first. You need C plus the eight items on this page. None of days 1 to 30 uses a template, a class or an exception, and a course that sends you away for three months is one you do not come back to.

Go deeper

Next

Day 1 writes the first kernel, and it prints nothing, which is the correct behaviour and catches almost everyone. The reason is the launch: the CPU keeps running after <<<>>> and the process can exit before the GPU has flushed a single character. Day 4 then turns the pointer arithmetic from item 5 into thread indexing, and day 11 is where the gap between consecutive and scattered addresses turns out to be worth 25 times the bandwidth on a Tesla T4.