What do global, device and host mean?
__global__ marks a kernel the host launches, __device__ marks a function only device code can call, and __host__ marks ordinary CPU code.
Those three are the answer everybody gives, and the answer everybody gives is out of date. The C++ language extensions appendix of the CUDA 13.3 guide lists five specifiers, not three: __host__, __device__, __tile__, __global__ and __tile_global__, in a table with separate "executed in" and "callable from" columns for host, SIMT and tile contexts. The Stack Overflow answer this query lands on has 118,614 views and dates from 2012, so it covers three of the five and none of the tile ones. A function with no specifier at all is __host__, which is why an innocent helper stops compiling the moment a kernel calls it.
The useful combination is __host__ __device__, which asks nvcc to compile the same function twice, once for each side. The two passes are not identical: __CUDA_ARCH__ is defined only in the device pass, and its value is the compute capability being compiled for, written as a three-digit xy0. That macro is how one body serves both sides, with an intrinsic on the device branch and plain C++ on the host branch. It is also how you check what your binary was built for, which is what day 3 does.
The failure is one direction of the arrow. Call a __host__ function from device code and nvcc stops with calling a __host__ function from a __device__ function is not allowed, usually because a kernel touched std::vector, std::string or an unannotated helper of your own. Device code gets a small subset of the standard library, and the subset is cuda::std from libcu++, not std. The other direction fails too: the host cannot call a __device__ function, so a helper you want on both sides needs both specifiers. See calling a host function from a device function for the ranked causes.
Measured
On a Tesla T4 (driver 595.84, CUDA 12.6, built with nvcc -O3 -arch=sm_75), day 3 launched a __global__ kernel that writes __CUDA_ARCH__ into device memory and copies it back. The value is 750, from a card that reports compute capability 7.5.
That is the split, measured. The host pass of the same file never sees 750, because __CUDA_ARCH__ does not exist there, and the host and the device halves of one translation unit disagree about what is defined. Note the version gap: the five-specifier table is CUDA 13.3 documentation, while this node runs toolkit 12.6, because its driver caps at CUDA 13.2. Captured 2026-08-30; transcript in code/day03-setup/evidence/run-2026-08-30.txt.
Diagram
CUDA 13.3 Programming Guide Table 39 separates execution context from caller context.
Code
A slice of code/day03-setup/devicequery.cu, which is a compiled target in this repo.
__global__ void reportCompiledArch(int* out, size_t n) {
const size_t i = blockIdx.x * static_cast<size_t>(blockDim.x) + threadIdx.x;
if (i < n) {
#if defined(__CUDA_ARCH__)
out[i] = __CUDA_ARCH__;
#else
out[i] = 0;
#endif
}
}
The #else branch is dead in this kernel and deliberate. Put the same guard in a __host__ __device__ function and both branches compile, one per pass, which is the whole reason the macro exists.
Related terms
Where you meet this
- Day 1, your first CUDA kernel, which owns this term and writes its first
__global__ - Day 3, how to set up CUDA, which owns the measurement above
- Day 5, vector addition, where a device helper first earns its specifier
- calling a host function from a device function
Sources
- CUDA Programming Guide, C++ language extensions, the function execution space specifiers and their callable-from table: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/cpp-language-extensions.html (checked 2026-08-29)
- The 2012 question this entry updates, 118,614 views: https://stackoverflow.com/questions/12373940/difference-between-global-and-device-functions (checked 2026-08-29)
Byline
Written by: unassigned. Reviewed by: unassigned. This entry is a draft and cannot publish until both are named people, and two different ones. Written on: not set. Last checked: not set. Numbers captured 2026-08-30 on the project's verification node.