What are the host and device in CUDA?
Host is the CPU and its memory, device is the GPU and its memory, and nothing crosses between them without a copy or a managed pointer.
Every CUDA program is two programs in one file. nvcc compiles the host code for your CPU and the device code for the GPU, and the two halves address different memory. A pointer from cudaMalloc is a number that means something to the GPU and nothing to your process, so reading d_a[0] on the host is not a slow read, it is a segmentation fault. That is why every pointer in this course carries h_ or d_: the prefix is the only thing standing between you and passing one where the other belongs. Which side a function runs on is set by its execution space specifier.
The two sides also run at the same time. A kernel launch returns to the host thread before the device has finished, and possibly before it has started, so the host carries on into its next statement while the GPU works. Nothing makes the host wait unless you write it: cudaDeviceSynchronize() for all outstanding work, a blocking cudaMemcpy, or a stream synchronize. A host-side stopwatch around a launch therefore measures the launch, not the kernel, which is the first thing day 9 takes apart.
Data crosses in one of two ways, and both cost. An explicit cudaMemcpy moves bytes over PCIe, which is roughly an order of magnitude slower than the GPU reading its own global memory. Unified memory gives you one pointer valid on both sides and lets the driver fault pages across on demand, which is convenient and not free. There is a third cost nobody writes: the first CUDA call in a process builds a context before your code does anything, and day 9 is where that gets a number rather than a shrug.
Measured
On a Tesla T4 (driver 595.84, CUDA 12.6, built with nvcc -O3 -arch=sm_75), day 1 printed from the host either side of one launch. The host's launch returned line came out before all eight device lines, and its sync returned line came out after them. The host got to its next statement while the kernel was still queued, and only the synchronize made it wait.
The two memories are separate on the same run. Day 3's device query on that card reports 14912 MiB of global memory and a 4096 KiB L2, none of which your host allocator can touch, and none of which knows anything about your std::vector. See how to set up CUDA for the full query output. Captured 2026-08-30; transcripts in code/day01-first-kernel/evidence/run-2026-08-30.txt and code/day03-setup/evidence/run-2026-08-30.txt.
Diagram
Alt text: two lanes on one time axis. The host lane runs launch, then its own printf, then blocks at the synchronize. The device lane starts after the launch and finishes during the host's block, so the host's second line always lands after every device line.
Related terms
Where you meet this
- Day 1, your first CUDA kernel, which owns the measurement above
- Day 3, how to set up CUDA, where the device query prints what the card has
- Day 5, vector addition, the first program that allocates and copies both ways
- Day 9, why your GPU looks slower than your CPU, where the crossing gets timed
Sources
- CUDA Programming Guide 2.1, "Intro to CUDA C++", on host and device memory and on launches continuing before the kernel completes: https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/intro-to-cuda-cpp.html (checked 2026-08-29)
- CUDA Programming Guide, C++ language extensions, on the specifiers that decide which side a function is compiled for: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/cpp-language-extensions.html (checked 2026-08-29)
Byline
Written by: unassigned. Reviewed by: unassigned. This entry is a draft and cannot publish until both are named people, and two different ones. Written on: not set. Last checked: not set. Numbers captured 2026-08-30 on the project's verification node.