CUDA ERROR
calling a __host__ function from a __device__ function
An unannotated function is host-only, so a kernel cannot call it. The CUDA 12.6 capture names __global__, not __device__, and raises two errors.
A function with no CUDA annotation is host-only, so device code has no compiled version of it to call.
| Kind | Compile-time, language rules. No cudaError code |
| Captured with | CUDA 12.6 (V12.6.85), 2026-09-01 |
| Source | code/errors/_toolchain/evidence/run-2026-09-01.txt |
Strings a reader might paste
Captured on this project's own node, from code/errors/_toolchain/evidence/run-2026-09-01.txt:
$ nvcc -std=c++17 -arch=sm_75 -c hostcall.cu
hostcall.cu(4): error: calling a __host__ function("tally(int)") from a __global__ function("k") is not allowed
hostcall.cu(4): error: identifier "tally" is undefined in device code
2 errors detected in the compilation of "hostcall.cu".
Two details from that capture that the folklore version of this error gets wrong. First, the message names __global__, not __device__, because the call was made directly from a kernel; the __device__ wording appears when the call is one level deeper, inside a device helper. Search engines see them as different strings. Second, you get two errors from one mistake, and the second one, identifier "tally" is undefined in device code, is the same bug wearing a disguise. People paste only the second line and go looking for a missing header.
The source that produced it is as plain as it gets: an ordinary static int tally(int x) with no annotation, and a kernel k that calls it.
Cause 1: your own helper has no annotation
The measured case. A function you wrote, with no __device__ on it, defaults to host-only.
static int tally(int x) { return x + 1; }
__global__ void k(int* out) { out[0] = tally(41); } // error
Fix: mark it __host__ __device__ so both compilations get a copy. Use __device__ alone only when the host genuinely must not call it.
Cause 2: the C++ standard library inside a kernel
std::vector, std::string, std::map, or anything from <iostream>. These are host functions, so the same rule applies, and the error often names a deeply mangled member you never wrote.
Fix: pass a raw pointer and a length instead. Device code gets a small subset of the standard library, and that subset is cuda::std from libcu++, not std.
Cause 3: a constructor or destructor you did not know was running
A class member whose constructor is host-only, instantiated in device code. Nothing in your source looks like a call, and the compiler names one anyway.
Fix: give the type __device__ constructors, or use a plain struct of plain data.
Confirm it
No tool. Read the first error, not the second, and look at the annotation on the function it names. The whole diagnosis is the annotation table:
| Annotation | Callable from host | Callable from device | Launched with <<<>>> |
|---|---|---|---|
none, or __host__ |
yes | no | no |
__device__ |
no | yes | no |
__host__ __device__ |
yes | yes | no |
__global__ |
launched, not called | (dynamic parallelism only) | yes |
If the function you need is in the "no" column for device, that is your answer and no build flag will change it.
Prevention
- Annotate small utility functions
__host__ __device__when you write them, not when the compiler asks. - Keep host containers out of anything a kernel touches. A struct crossing to the device holds plain data and raw pointers, nothing that owns memory.
Related errors
a value of type "void *" cannot be assigned: the other day-one language rule, from the same capture.invalid device function(98): the run-time relative, where the symbol compiled but is missing on this device.invalid device symbol(13): the same registration table, for variables.
The lesson
Day 0 introduces the annotations; day 39 resolves the standard-library half by introducing Thrust, CUB and libcu++, which is where cuda::std comes from. The full code table is on the error hub; toolkit setup is /setup/install-cuda.
Written 2026-09-01. Compiler output from code/errors/_toolchain/evidence/run-2026-09-01.txt, captured 2026-09-01 with CUDA 12.6 (V12.6.85) on ornn-internal-t4-spot-5v08. This is a compile-time error, so no GPU was needed to reproduce it. Author and reviewer: not yet assigned; this page does not publish until both are named.