CUDA errorsdraft
Reproduced 2026-09-01

CUDA ERROR

calling a __host__ function from a __device__ function

An unannotated function is host-only, so a kernel cannot call it. The CUDA 12.6 capture names __global__, not __device__, and raises two errors.

A function with no CUDA annotation is host-only, so device code has no compiled version of it to call.

Kind Compile-time, language rules. No cudaError code
Captured with CUDA 12.6 (V12.6.85), 2026-09-01
Source code/errors/_toolchain/evidence/run-2026-09-01.txt

Strings a reader might paste

Captured on this project's own node, from code/errors/_toolchain/evidence/run-2026-09-01.txt:

$ nvcc -std=c++17 -arch=sm_75 -c hostcall.cu
hostcall.cu(4): error: calling a __host__ function("tally(int)") from a __global__ function("k") is not allowed

hostcall.cu(4): error: identifier "tally" is undefined in device code

2 errors detected in the compilation of "hostcall.cu".

Two details from that capture that the folklore version of this error gets wrong. First, the message names __global__, not __device__, because the call was made directly from a kernel; the __device__ wording appears when the call is one level deeper, inside a device helper. Search engines see them as different strings. Second, you get two errors from one mistake, and the second one, identifier "tally" is undefined in device code, is the same bug wearing a disguise. People paste only the second line and go looking for a missing header.

The source that produced it is as plain as it gets: an ordinary static int tally(int x) with no annotation, and a kernel k that calls it.

Cause 1: your own helper has no annotation

The measured case. A function you wrote, with no __device__ on it, defaults to host-only.

static int tally(int x) { return x + 1; }
__global__ void k(int* out) { out[0] = tally(41); }   // error

Fix: mark it __host__ __device__ so both compilations get a copy. Use __device__ alone only when the host genuinely must not call it.

Cause 2: the C++ standard library inside a kernel

std::vector, std::string, std::map, or anything from <iostream>. These are host functions, so the same rule applies, and the error often names a deeply mangled member you never wrote.

Fix: pass a raw pointer and a length instead. Device code gets a small subset of the standard library, and that subset is cuda::std from libcu++, not std.

Cause 3: a constructor or destructor you did not know was running

A class member whose constructor is host-only, instantiated in device code. Nothing in your source looks like a call, and the compiler names one anyway.

Fix: give the type __device__ constructors, or use a plain struct of plain data.

Confirm it

No tool. Read the first error, not the second, and look at the annotation on the function it names. The whole diagnosis is the annotation table:

Annotation Callable from host Callable from device Launched with <<<>>>
none, or __host__ yes no no
__device__ no yes no
__host__ __device__ yes yes no
__global__ launched, not called (dynamic parallelism only) yes

If the function you need is in the "no" column for device, that is your answer and no build flag will change it.

Prevention

  • Annotate small utility functions __host__ __device__ when you write them, not when the compiler asks.
  • Keep host containers out of anything a kernel touches. A struct crossing to the device holds plain data and raw pointers, nothing that owns memory.

Related errors

The lesson

Day 0 introduces the annotations; day 39 resolves the standard-library half by introducing Thrust, CUB and libcu++, which is where cuda::std comes from. The full code table is on the error hub; toolkit setup is /setup/install-cuda.


Written 2026-09-01. Compiler output from code/errors/_toolchain/evidence/run-2026-09-01.txt, captured 2026-09-01 with CUDA 12.6 (V12.6.85) on ornn-internal-t4-spot-5v08. This is a compile-time error, so no GPU was needed to reproduce it. Author and reviewer: not yet assigned; this page does not publish until both are named.